@dailephd/my-frontend-observer 0.9.1 → 0.10.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (150) hide show
  1. package/CHANGELOG.md +490 -471
  2. package/LICENSE +21 -21
  3. package/README.md +375 -357
  4. package/dist/application/projectCheckService.d.ts +6 -0
  5. package/dist/application/projectCheckService.js +8 -1
  6. package/dist/application/projectCheckService.js.map +1 -1
  7. package/dist/application/projectWorkflowService.d.ts +7 -2
  8. package/dist/application/projectWorkflowService.js +10 -3
  9. package/dist/application/projectWorkflowService.js.map +1 -1
  10. package/dist/application/visualChangeAgentHandoffService.d.ts +28 -0
  11. package/dist/application/visualChangeAgentHandoffService.js +111 -0
  12. package/dist/application/visualChangeAgentHandoffService.js.map +1 -0
  13. package/dist/application/visualChangeProjectWorkflowService.d.ts +95 -0
  14. package/dist/application/visualChangeProjectWorkflowService.js +376 -0
  15. package/dist/application/visualChangeProjectWorkflowService.js.map +1 -0
  16. package/dist/application/visualChangeReviewService.d.ts +50 -0
  17. package/dist/application/visualChangeReviewService.js +69 -0
  18. package/dist/application/visualChangeReviewService.js.map +1 -0
  19. package/dist/application/visualChangeWorkflowPersistenceService.d.ts +26 -0
  20. package/dist/application/visualChangeWorkflowPersistenceService.js +15 -0
  21. package/dist/application/visualChangeWorkflowPersistenceService.js.map +1 -0
  22. package/dist/artifacts/visualChangeWorkflowArtifactReader.d.ts +9 -0
  23. package/dist/artifacts/visualChangeWorkflowArtifactReader.js +47 -0
  24. package/dist/artifacts/visualChangeWorkflowArtifactReader.js.map +1 -0
  25. package/dist/artifacts/visualChangeWorkflowArtifactWriter.d.ts +20 -0
  26. package/dist/artifacts/visualChangeWorkflowArtifactWriter.js +41 -0
  27. package/dist/artifacts/visualChangeWorkflowArtifactWriter.js.map +1 -0
  28. package/dist/cli.js +9 -7
  29. package/dist/cli.js.map +1 -1
  30. package/dist/domain/visualChangeAgentHandoff.d.ts +82 -0
  31. package/dist/domain/visualChangeAgentHandoff.js +80 -0
  32. package/dist/domain/visualChangeAgentHandoff.js.map +1 -0
  33. package/dist/domain/visualChangeAgentHandoffSerialization.d.ts +2 -0
  34. package/dist/domain/visualChangeAgentHandoffSerialization.js +11 -0
  35. package/dist/domain/visualChangeAgentHandoffSerialization.js.map +1 -0
  36. package/dist/domain/visualChangeCycle.d.ts +8 -0
  37. package/dist/domain/visualChangeCycle.js +7 -0
  38. package/dist/domain/visualChangeCycle.js.map +1 -0
  39. package/dist/domain/visualChangeWorkflow.d.ts +125 -0
  40. package/dist/domain/visualChangeWorkflow.js +109 -0
  41. package/dist/domain/visualChangeWorkflow.js.map +1 -0
  42. package/dist/domain/visualChangeWorkflowIdentity.d.ts +5 -0
  43. package/dist/domain/visualChangeWorkflowIdentity.js +24 -0
  44. package/dist/domain/visualChangeWorkflowIdentity.js.map +1 -0
  45. package/dist/index.d.ts +21 -1
  46. package/dist/index.js +12 -1
  47. package/dist/index.js.map +1 -1
  48. package/dist/projectWorkflow/projectPaths.d.ts +3 -0
  49. package/dist/projectWorkflow/projectPaths.js +7 -0
  50. package/dist/projectWorkflow/projectPaths.js.map +1 -1
  51. package/dist/viewer/assets/index-DglJ6f28.css +1 -0
  52. package/dist/viewer/assets/index-DsODREY5.js +9 -0
  53. package/dist/viewer/index.html +15 -15
  54. package/dist/viewer/sw.js +1 -1
  55. package/dist/viewerServer/evidence/classify.d.ts +3 -1
  56. package/dist/viewerServer/evidence/classify.js +10 -0
  57. package/dist/viewerServer/evidence/classify.js.map +1 -1
  58. package/dist/viewerServer/evidence/handles.js +1 -0
  59. package/dist/viewerServer/evidence/handles.js.map +1 -1
  60. package/dist/viewerServer/evidence/projection.d.ts +5 -0
  61. package/dist/viewerServer/evidence/projection.js +19 -0
  62. package/dist/viewerServer/evidence/projection.js.map +1 -1
  63. package/dist/viewerServer/evidence/visualChangeWorkflowView.d.ts +31 -0
  64. package/dist/viewerServer/evidence/visualChangeWorkflowView.js +36 -0
  65. package/dist/viewerServer/evidence/visualChangeWorkflowView.js.map +1 -0
  66. package/dist/viewerServer/httpServer.js +323 -1
  67. package/dist/viewerServer/httpServer.js.map +1 -1
  68. package/dist/viewerServer/referenceApproval.d.ts +22 -0
  69. package/dist/viewerServer/referenceApproval.js +42 -0
  70. package/dist/viewerServer/referenceApproval.js.map +1 -0
  71. package/dist/viewerServer/referenceVisualChangeAuthoring.d.ts +28 -0
  72. package/dist/viewerServer/referenceVisualChangeAuthoring.js +134 -0
  73. package/dist/viewerServer/referenceVisualChangeAuthoring.js.map +1 -0
  74. package/dist/viewerServer/runtimeVisualChangeAuthoring.d.ts +33 -0
  75. package/dist/viewerServer/runtimeVisualChangeAuthoring.js +81 -0
  76. package/dist/viewerServer/runtimeVisualChangeAuthoring.js.map +1 -0
  77. package/dist/viewerServer/visualChangeAuthoring.d.ts +46 -0
  78. package/dist/viewerServer/visualChangeAuthoring.js +63 -0
  79. package/dist/viewerServer/visualChangeAuthoring.js.map +1 -0
  80. package/dist/viewerServer/visualChangeHandoff.d.ts +23 -0
  81. package/dist/viewerServer/visualChangeHandoff.js +31 -0
  82. package/dist/viewerServer/visualChangeHandoff.js.map +1 -0
  83. package/dist/viewerServer/visualChangeReview.d.ts +30 -0
  84. package/dist/viewerServer/visualChangeReview.js +46 -0
  85. package/dist/viewerServer/visualChangeReview.js.map +1 -0
  86. package/docs/ARCHITECTURE.md +1394 -1373
  87. package/docs/CI_CD.md +349 -327
  88. package/docs/COMMANDS.md +1035 -1012
  89. package/docs/CONTRACTS.md +1971 -1926
  90. package/docs/CURRENT_STATE.md +1277 -1238
  91. package/docs/DEVELOPMENT.md +240 -237
  92. package/docs/DOCUMENTATION_PRESERVATION_POLICY.md +50 -50
  93. package/docs/PROJECT_DESCRIPTION.md +2248 -2224
  94. package/docs/PROJECT_MILESTONES.md +2681 -2558
  95. package/docs/PROJECT_OVERVIEW.md +200 -191
  96. package/docs/QUICKSTART.md +100 -96
  97. package/docs/RELEASE.md +37 -33
  98. package/docs/ROADMAP.md +1105 -1033
  99. package/docs/SECURITY.md +297 -275
  100. package/docs/WORKFLOWS.md +806 -770
  101. package/docs/plans/v0.10-implementation-plan.md +1509 -0
  102. package/docs/plans/v0.8-implementation-plan.md +655 -655
  103. package/docs/plans/v0.8.1-cli-usability-patch-plan.md +505 -505
  104. package/docs/plans/v0.9-implementation-plan.md +1529 -1529
  105. package/docs/plans/v0.9.1-implementation-plan.md +468 -468
  106. package/docs/reports/v0.10-batch1-visual-change-workflow-foundation.md +102 -0
  107. package/docs/reports/v0.10-batch2-project-composition-check-recording.md +103 -0
  108. package/docs/reports/v0.10-batch3-viewer-visual-change-workspace.md +93 -0
  109. package/docs/reports/v0.10-batch4-actual-frontend-entry.md +59 -0
  110. package/docs/reports/v0.10-batch5-reference-driven-entry.md +238 -0
  111. package/docs/reports/v0.10-batch6-coding-agent-handoff.md +85 -0
  112. package/docs/reports/v0.10-batch7-correction-review-acceptance.md +145 -0
  113. package/docs/reports/v0.10-batch8-integrated-acceptance.md +109 -0
  114. package/docs/reports/v0.10-implementation-completeness-documentation-reconciliation.md +344 -0
  115. package/docs/reports/v0.10-pre-release-readiness.md +120 -0
  116. package/docs/reports/v0.10-release-preparation.md +70 -0
  117. package/docs/reports/v0.10.1-project-check-baseline-context-implementation.md +86 -0
  118. package/docs/reports/v0.7-bounded-fidelity-context-prompt7.md +243 -243
  119. package/docs/reports/v0.7-implementation-completeness-documentation-reconciliation.md +497 -497
  120. package/docs/reports/v0.7-pre-release-readiness.md +337 -337
  121. package/docs/reports/v0.7-reference-binding-prompt5.md +223 -223
  122. package/docs/reports/v0.7-reference-compatibility-prompt4.md +234 -234
  123. package/docs/reports/v0.7-reference-correction-workflow-prompt8.md +222 -222
  124. package/docs/reports/v0.7-reference-fidelity-prompt6.md +216 -216
  125. package/docs/reports/v0.7-reference-foundation-prompt1.md +151 -151
  126. package/docs/reports/v0.7-reference-regions-prompt2.md +195 -195
  127. package/docs/reports/v0.7-reference-requirements-prompt3.md +217 -217
  128. package/docs/reports/v0.7-release-prep.md +423 -423
  129. package/docs/reports/v0.8-binding-fidelity-interaction-batch6.md +279 -279
  130. package/docs/reports/v0.8-bounded-context-correlation-batch7.md +233 -233
  131. package/docs/reports/v0.8-comparison-contract-inspection-batch4.md +279 -279
  132. package/docs/reports/v0.8-evidence-index-readers-batch2.md +247 -247
  133. package/docs/reports/v0.8-implementation-completeness-documentation-reconciliation.md +741 -741
  134. package/docs/reports/v0.8-integrated-viewer-acceptance-batch8.md +128 -128
  135. package/docs/reports/v0.8-observation-svg-inspection-batch3.md +223 -223
  136. package/docs/reports/v0.8-prerelease-readiness-cross-platform-security-code-rot.md +687 -687
  137. package/docs/reports/v0.8-reference-candidate-inspection-batch5.md +232 -232
  138. package/docs/reports/v0.8-viewer-runtime-pwa-batch1.md +278 -278
  139. package/docs/reports/v0.8.1-implementation-completeness-documentation-reconciliation.md +114 -114
  140. package/docs/reports/v0.8.1-prerelease-readiness-cross-platform-security-code-rot.md +170 -170
  141. package/docs/reports/v0.9-architecture-retrieval.md +14 -37
  142. package/docs/reports/v0.9-final-pre-release-readiness.md +209 -209
  143. package/docs/reports/v0.9-final-readiness-corrections.md +530 -530
  144. package/docs/reports/v0.9-pre-release-readiness.md +169 -169
  145. package/docs/reports/v0.9.1-batch1-pwa-hard-gate-isolation.md +359 -359
  146. package/docs/reports/v0.9.1-batch2-hard-gate-validation-integration.md +262 -262
  147. package/docs/reports/v0.9.1-pre-release-readiness.md +206 -206
  148. package/package.json +59 -59
  149. package/dist/viewer/assets/index-BN41MI7m.css +0 -1
  150. package/dist/viewer/assets/index-CkKXnlrI.js +0 -9
package/docs/CONTRACTS.md CHANGED
@@ -1,1926 +1,1971 @@
1
- # Contracts
2
-
3
- ## Current contracts
4
-
5
- The observation artifact contract is published in the current
6
- `my-frontend-observer@0.7.0` package and proven both from the source checkout
7
- and from the packed npm tarball, on Windows, Linux, and macOS. The observation
8
- schema is `1.2.0` (see "v0.2 target contract" and "v0.3 scroll scenario
9
- contract" below):
10
-
11
- - artifact kind `my-frontend-observer/observation`, schema version `1.2.0`
12
- (independent of the package version);
13
- - one artifact root per observation, `<outputLocation>/<observationId>/`,
14
- containing exactly `manifest.json` (the full `ObservationArtifact`, with
15
- page/target evidence embedded inline) and `screenshot.png` - there is no
16
- separate `evidence.json`;
17
- - `manifest.json` is written last, after `screenshot.png`, via one atomic
18
- directory rename, so a consumer never observes a partially-written
19
- artifact; a filesystem failure anywhere in that sequence reports the
20
- `artifact-write-failure` diagnostic and leaves no completed artifact;
21
- - internal artifact references (e.g. `screenshot.png`) are relative to the
22
- artifact root, never an absolute machine path; the observation's logical
23
- identity is its `observationId`, not its filesystem location;
24
- - evidence states `available`, `unavailable`, `not-applicable`, `partial`;
25
- evidence sources `browser`, `computed-browser`, `derived`;
26
- - a stable diagnostic vocabulary (`src/domain/diagnostics.ts`) and completion
27
- states `complete`, `partial`, `warning`, `invalid-request`, `fatal`
28
- (`src/domain/completion.ts`);
29
- - observation/request identity, producer/package identity, and browser
30
- provenance are all present in every persisted manifest.
31
-
32
- This contract is implemented and published; no public programmatic-API
33
- compatibility promise has been made for the observation engine itself. v0.6
34
- additionally publishes the bounded-agent-context/correlation programmatic surface
35
- described later in this document.
36
-
37
- ## v0.2 target contract (shipped as part of this release)
38
-
39
- v0.2 introduces a canonical target-configuration model: each configured target has a stable
40
- observer-level `name` plus an ordered array of bounded `locators`
41
- (`role`, `id`, `data-attribute`, `semantic-element`, `css`, `text`). This
42
- identity is distinct from both the browser locator definition that resolves
43
- it and any source-code identity. The legacy `{name, selector}` shape remains
44
- accepted and normalizes to a one-item `css` locator, so every published
45
- `0.1.0` CLI invocation continues to work unchanged. Locator precedence is the
46
- configured array order; resolution stops on the first unique match, on any
47
- ambiguous match (never falling through to a later locator), or on an
48
- unevaluable locator - never silently. All six frozen locator kinds are now
49
- resolved against a real Chromium page (`role` via Playwright's accessibility-
50
- role/name locator with exact name matching, `id`/`data-attribute` via exact
51
- CSS attribute-equals matching that never reinterprets the configured value as
52
- selector syntax, `semantic-element` via the frozen tag set, `css` via the
53
- existing v0.1 behavior, `text` via exact-text matching only); every kind
54
- converges on the same measurement path, so locator strategy never changes the
55
- resulting target evidence shape.
56
-
57
- Each resolved target's evidence record additionally carries three bounded
58
- fields: `semanticState` (a first family of `disabled`/`expanded`/
59
- `checked`/`selected`/`pressed`/`current` values read from the element's own
60
- native form-control properties and explicit `aria-*` attributes - a key is
61
- present only when the browser exposes that state as applicable to this
62
- element, so an explicit `false` is always distinguishable from "not
63
- applicable"; `not-applicable` when no supported state applies at all);
64
- `landmark` (derived only from the already-captured browser-exposed
65
- role - never from locator kind or HTML tag - against the standard landmark
66
- role set `banner`/`navigation`/`main`/`complementary`/`contentinfo`/`form`/
67
- `region`/`search`); and `containment` (bounded DOM containment checked only
68
- among the other explicitly configured targets in the same observation, in
69
- configured order, never a layout/relationship graph - `available` when every
70
- other configured target was itself resolved and checked, `partial` when one
71
- or more could not be, `unavailable` when the target itself never resolved).
72
- Stable observer target identity is proven, not just declared: the same
73
- target configuration produces the same `requestId` across repeated
74
- observations (with a fresh `observationId` each time); changing a target's
75
- locator strategy while keeping its stable name changes `requestId` but not
76
- the `targetEvidence` key; and actual runtime disappearance of a
77
- still-configured target changes only its resolution status, never the
78
- `requestId`.
79
-
80
- The full canonical semantic target model above is reachable through the
81
- real public CLI: `my-frontend-observer observe --targets-file <json-file>`
82
- supplies the structured `{ "targets": [...] }` collection (see
83
- `docs/COMMANDS.md` "Structured semantic targets") as an alternative to the
84
- existing `--target id=css-selector` shorthand - the two are mutually
85
- exclusive per invocation, and both converge on the same
86
- `normalizeRequest()`/browser-resolver/artifact path, so a semantic
87
- observation produces exactly the same `manifest.json` shape as a
88
- CSS-shorthand one. Schema `1.1.0` was the v0.2 published artifact schema;
89
- schema `1.2.0` has been emitted since v0.3 and remains the observation schema
90
- in the current published v0.7.0 package, for both target-input modes
91
- (target semantics are unchanged from v0.2 - see the v0.3 scroll scenario
92
- contract below for what schema `1.2.0` actually adds). `--targets-file`'s
93
- local input path is never part of the persisted request identity or
94
- artifact.
95
-
96
- ## v0.3 scroll scenario contract (shipped as part of this release)
97
-
98
- v0.3 introduces one optional, additive request/evidence concern: a bounded
99
- runtime scroll scenario, schema `1.2.0`.
100
-
101
- A normalized request may carry `scrollScenario: { action }` with exactly one
102
- of two frozen action kinds:
103
-
104
- - `{ "kind": "window-scroll-by", "deltaX": <int>, "deltaY": <int> }`
105
- - `{ "kind": "target-scroll-by", "target": "<stable target name>", "deltaX": <int>, "deltaY": <int> }`
106
-
107
- `deltaX`/`deltaY` are signed integers bounded to `[-20000, 20000]`; at least
108
- one must be non-zero. `target-scroll-by.target` refers only to an existing
109
- stable configured target `name` (never a selector) and resolves through the
110
- same canonical `resolveConfiguredTargets` algorithm every v0.2 locator kind
111
- already uses - there is no second target-resolution path. A request with no
112
- scenario normalizes and identifies exactly as it did before v0.3.
113
-
114
- Execution (both action kinds share one code path): perform the immediate,
115
- non-smooth scroll (`window.scrollBy`/`element.scrollBy`, `behavior:
116
- 'instant'`) on the already-navigated, already-ready page; wait exactly two
117
- `requestAnimationFrame` cycles; capture a final runtime snapshot. No second
118
- browser, page, or navigation is ever created. The resulting scroll position
119
- is browser-authoritative and may be clamped by document/element boundaries;
120
- a scenario producing no movement is still a valid, successfully persisted
121
- observation.
122
-
123
- The scenario evidence lives entirely inside the existing `manifest.json` as
124
- one additional optional `scrollScenarioEvidence` field on `ObservationArtifact`
125
- - there is no separate `scroll.json`/`scenario.json`. It contains:
126
-
127
- - `initial`/`final`: bounded `ScrollRuntimeSnapshot`s (window `scrollX`/
128
- `scrollY`; the browser's own scrolling-root/`documentElement`/`body`
129
- metrics; per-configured-target `scrollTop`/`scrollLeft`/`scrollWidth`/
130
- `scrollHeight`/`clientWidth`/`clientHeight`, actual overflow, bounding
131
- rectangle, and viewport relation);
132
- - `transition`: bounded before/after change evidence (window scroll deltas;
133
- per-target `scrollTop`/`scrollLeft`/bounding-position/viewport-relation
134
- changes; `enteredViewport`/`leftViewport`) - never a generic recursive
135
- diff, and a target is simply omitted when either side's evidence isn't
136
- itself usable (e.g. it never resolved);
137
- - `scrollOwner`: one derived `EvidenceField<ScrollOwnerInterpretation>`
138
- (`document` | `target:<stable-name>` | `none` | `indeterminate`), always
139
- `source: "derived"` with non-empty `derivedFrom` naming the exact
140
- contributing scroll-position measurements. Ownership is derived only from
141
- observed `scrollTop`/`scrollLeft`/`window.scrollX`/`window.scrollY`
142
- changes - never from bounding-rectangle movement (which moves for every
143
- configured target whenever the document scrolls), computed overflow,
144
- `position: fixed`/`sticky`, or DOM hierarchy.
145
-
146
- Actual dimensional overflow (`scrollWidth > clientWidth` /
147
- `scrollHeight > clientHeight`) is always reported separately from the
148
- computed `overflow-x`/`overflow-y` CSS declaration; a declared
149
- `overflow: auto` container with content that fits produces
150
- `horizontalOverflow`/`verticalOverflow: false`. Viewport relation
151
- (`above`/`intersecting`/`below`, `intersectsViewport`, `fullyWithinViewport`)
152
- is derived only from bounding geometry plus viewport size, relative to the
153
- browser viewport; a hidden/non-rendered target's viewport relation is
154
- `not-applicable`, never a fabricated geometry claim - hidden and offscreen
155
- remain distinct evidence concepts, and the existing `target-hidden`
156
- diagnostic is unaffected.
157
-
158
- The ordinary, already-existing `pageEvidence`/`targetEvidence`/
159
- `screenshot.png` for a scenario observation always describe this same final
160
- post-action state, never the pre-action state.
161
-
162
- The scenario request participates in `requestId`; the runtime result
163
- (actual scroll distance, clamping, or scroll-owner outcome) never does. The
164
- public entry point is `my-frontend-observer observe --scroll-scenario-file
165
- <json-file>` (see `docs/COMMANDS.md`); the file supplies the scenario value
166
- directly, and its local path is operational input only, exactly like
167
- `--targets-file`'s path - never persisted, never part of request identity.
168
-
169
- ## v0.4 comparison contract (shipped as part of this release)
170
-
171
- **Current status: shipped as part of the published `my-frontend-observer@0.4.0`
172
- package and unchanged through the current `0.7.0` release.** Observation
173
- schema remains `1.2.0`. Comparison is a distinct artifact kind and schema,
174
- never a bump to the observation schema:
175
-
176
- - artifact kind: `my-frontend-observer/comparison`;
177
- - comparison schema: `1.0.0`.
178
-
179
- **Geometry tolerance**: `ComparisonConfig.geometryTolerancePx`, default
180
- `0.5` CSS px, bounded `[0, 10]`. Suppresses insignificant subpixel noise
181
- only - never a design contract, never permission for a change.
182
-
183
- **Layout relationship graph**: `deriveLayoutRelationships(observation,
184
- options?)` derives, per observation, a bounded `LayoutRelationshipGraph`
185
- among configured targets only (≤20 targets, ≤190 unordered pairs):
186
- horizontal order (`left-of`/`right-of`/`horizontally-overlapping`),
187
- vertical order (`above`/`below`/`vertically-overlapping`), area overlap
188
- (`overlaps`/`does-not-overlap`), relative width (`wider-than`/
189
- `narrower-than`/`equal-width-within-tolerance`), geometric fit
190
- (`fits-inside`/`does-not-fit-inside` - geometry-only, deliberately distinct
191
- from DOM containment), vertical sequencing (`follows-vertically`), and one
192
- page-level relationship (`document-width-fits-viewport`/
193
- `document-width-exceeds-viewport`). Every relationship carries explicit
194
- evidence-path provenance back to the source observation. A configured
195
- target lacking usable geometry is listed as honestly unresolved
196
- (`not-found`/`ambiguous`/`unavailable`/`hidden`), never fabricated as a
197
- zero-sized region.
198
-
199
- **Comparability**: evaluated before any rendered difference, using exactly
200
- three states (`comparable`/`comparable-with-warnings`/`incomparable`) with
201
- structured reasons, never a bare boolean. Hard incompatibilities (page URL,
202
- viewport, browser engine, scroll-scenario configuration mismatch) force
203
- `incomparable`; producer-version, browser-version, and target-configuration
204
- differences are warning-only; theme/authenticated-state/application-state
205
- identity are recorded as `unassessed` dimensions the observer does not yet
206
- model - never silently claimed identical. An `incomparable` result still
207
- persists a structurally valid `ComparisonArtifact` with empty rendered
208
- differences, not a fabricated comparison.
209
-
210
- **Difference categories**: `appeared`/`disappeared` (only for a stable
211
- target name configured on both sides, transitioning between a definite
212
- `not-found` and `matched` resolution status - never for a target merely
213
- added/removed from configuration, which is its own separate
214
- `configurationChanges` entry), `moved`/`resized` (tolerance-aware, a target
215
- may be both), `visibility-changed`, `clipping-changed` (reusing the
216
- canonical `deriveTargetClipping` helper, never re-derived), `horizontal-
217
- overflow-changed`/`vertical-overflow-changed` (actual dimensional overflow,
218
- reusing the existing `deriveOverflowEvidence` helper - never inferred from
219
- a CSS declaration alone), `containment-changed` (reusing existing v0.2
220
- `TargetContainment` evidence), `page-size-changed`, `scroll-owner-changed`
221
- (comparing `scrollScenarioEvidence.scrollOwner` only when scenario
222
- *configuration* already matched), `relative-position-changed` (a relation
223
- in the horizontal-order/vertical-order/area-overlap families changed - kept
224
- distinct from plain absolute target movement) and `relationship-changed`
225
- (every other relationship-family transition). Relationship changes are
226
- matched by structural identity (family + subject/related target, or the
227
- page-level key), never by array position.
228
-
229
- **Explicit dependency evidence**: `ComparisonConfig.expectedDependencies`
230
- lets a caller declare an expected relationship between two targets' numeric
231
- properties (`x`/`y`/`width`/`height`) and directions (`increase`/
232
- `decrease`/`change`/`unchanged`), always carrying `source:
233
- "explicit-config"`. The observer never synthesizes a declaration from
234
- observed co-change. Each declaration evaluates independently to exactly one
235
- of `consistent`/`not-observed`/`contradictory-to-declaration`/
236
- `unavailable` - never a causal claim (no `causedBy`/`causalConfidence`/
237
- `causalScore`/`dependencyStrength`) and never a PASS/FAIL/approval verdict.
238
- That distinction (evidence vs. contract verdict) is the boundary between
239
- v0.4 and v0.5+.
240
-
241
- **Comparison identity**: `comparisonRequestId` is a pure, deterministic
242
- function of `{beforeObservationId, afterObservationId, normalized
243
- ComparisonConfig}` - direction-sensitive (`compare(A, B) !==
244
- compare(B, A)`), and never includes an operational filesystem path.
245
- `comparisonId` is fresh per execution (same pattern as `observationId`).
246
-
247
- **Source references**: the comparison artifact retains enough logical
248
- identity to trace back to its authoritative source observations
249
- (`observationId`, `requestId`, `producer`, `observationSchemaVersion`, and
250
- the source `screenshot.path`) without embedding the full
251
- `ObservationArtifact` or copying screenshot bytes. The persisted comparison
252
- directory contains `manifest.json` only.
253
-
254
- The public entry point is `my-frontend-observer compare --before <root>
255
- --after <root> --output <directory> [--config-file <json-file>]` (see
256
- `docs/COMMANDS.md`) - comparison itself never launches a browser.
257
-
258
- ## v0.5 frontend contract and evaluation (shipped as part of this release)
259
-
260
- Downstream of the v0.4 observation/comparison/relationship evidence above,
261
- `src/domain/frontendContracts.ts` freezes the v0.5 contract/change-scope
262
- model, `src/domain/frontendContractIdentity.ts` freezes deterministic
263
- contract/baseline/clause identity, and `src/domain/frontendContractEvaluation.ts`
264
- implements the one canonical pure evaluation engine. Baseline/per-change
265
- contract persistence, evaluation-artifact persistence, explicit baseline
266
- approval, and public CLI exposure are all implemented and shipped (see
267
- "v0.5 contract and evaluation persistence" and "v0.5 public contract/
268
- evaluation commands" below).
269
-
270
- **Contract classes**: a `PersistentBaselineContract` (append/supersession-based
271
- history via an optional `supersedesBaselineId`) and a `PerChangeContract`
272
- (the allowed scope of one requested change). Both share `artifactKind:
273
- "my-frontend-observer/frontend-contract"` and `schemaVersion: "1.0.0"` - an
274
- independent family from the observation (`1.2.0`) and comparison (`1.0.0`)
275
- schemas; the frontend-contract schema constant happens to share the version
276
- string `1.0.0` with comparison's by coincidence only.
277
-
278
- **Four authored categories, one derived classification**: every per-change
279
- clause is authored as exactly one of `requested`, `expected-dependent`,
280
- `protected`, or `preserved`. `unexpected` is a fifth, *derived-only*
281
- classification the evaluator produces for a meaningful rendered difference no
282
- active clause accounts for - it can never be authored as a permission.
283
-
284
- **Bounded contract primitives**: 15 frozen `ContractPrimitive` kinds cover
285
- visibility, clipping, width bounds, non-overlap, relative width, vertical
286
- sequence, geometric fit (explicitly distinct from DOM containment),
287
- document-width-vs-viewport, scroll ownership, initial-viewport position,
288
- relationship-unchanged, and property-unchanged/increases/decreases - a closed
289
- vocabulary, never a generic expression language.
290
-
291
- **Contract tolerance**: `exact` / `absolute-px` / `percent`, independent of
292
- `ComparisonConfig.geometryTolerancePx` (which only suppresses insignificant
293
- comparison noise and is never contract authorization). Percent tolerance's
294
- denominator is the absolute before-value.
295
-
296
- **Required vs. permitted expected-dependent**: `required` clauses must occur
297
- compliantly to pass; `permitted` clauses accept no change or a compliant
298
- change, and fail only on a strictly contradictory change.
299
-
300
- **Evaluation result vocabulary**: each clause resolves to `pass` / `fail` /
301
- `unavailable` (with a required non-empty reason - required evidence gaps and
302
- an `incomparable` source comparison never fabricate a `pass`) / `conflict`
303
- (with at least two `conflictingClauseIds` - covers both an unresolved
304
- baseline/per-change contradiction and an unknown `supersedesBaselineClauseIds`
305
- reference). The overall verdict is `PASS` only when every clause result is
306
- `pass` and no unexpected change remains; otherwise `FAIL` - there is no
307
- partial-pass scoring.
308
-
309
- **Explicit supersession, never inferred**: a per-change clause may list
310
- `supersedesBaselineClauseIds` to remove specific baseline clauses from active
311
- evaluation. Two clauses that structurally contradict each other on the same
312
- (target, property) without explicit supersession produce a `conflict`, never
313
- a silent preference for one side.
314
-
315
- **Reuses existing v0.4 evidence directly**: the evaluator consumes an
316
- already-computed `ComparisonArtifact` (`differences`, `relationshipChanges`,
317
- `relationshipsBefore`/`relationshipsAfter`, `comparability`) and the source
318
- `ObservationArtifact` pair - it never re-launches a browser, re-resolves a
319
- target, or reimplements clipping/relationship/scroll-owner derivation.
320
- Unexpected-change derivation reads `ComparisonArtifact.differences` only
321
- (which already includes one difference per relationship change), so a single
322
- logical transition is never double-counted.
323
-
324
- ## v0.5 contract and evaluation persistence (shipped as part of this release)
325
-
326
- Persistence consumes the frozen v0.5 domain above; it never redefines it.
327
- `src/artifacts/frontendContractArtifactWriter.ts`/`frontendContractArtifactReader.ts`
328
- persist and read both `PersistentBaselineContract` and `PerChangeContract`
329
- symmetrically (both already share `CONTRACT_ARTIFACT_KIND`/`CONTRACT_SCHEMA_VERSION`,
330
- so one writer/reader pair serves both contract classes) as
331
- `<outputLocation>/<baselineId|contractId>/manifest.json`, following the same
332
- atomic-write discipline as `artifacts/artifactWriter.ts`/`artifacts/comparisonArtifactWriter.ts`
333
- (sibling temporary directory, then one atomic rename; an existing directory at
334
- the final identity is a genuine collision and is rejected, never overwritten -
335
- prior baseline history is never rewritten). `src/artifacts/comparisonArtifactReader.ts`
336
- is a new Batch 3 addition (no comparison reader existed before) mirroring
337
- `artifacts/artifactReader.ts`'s discipline exactly, changing no comparison
338
- semantics and keeping comparison schema `1.0.0`.
339
-
340
- **Evaluation artifact envelope**: Batch 1 froze the evaluation-result
341
- vocabulary (`ClauseEvaluationResult`, `OverallVerdict`) but not a persistable
342
- envelope, so `src/domain/frontendContractEvaluationArtifact.ts` adds exactly
343
- that - `artifactKind: "my-frontend-observer/frontend-contract-evaluation"`,
344
- `schemaVersion: "1.0.0"` (its own independent family, distinct from
345
- observation/comparison/frontend-contract), an `evaluationId`/`evaluationRequestId`
346
- pair, bounded `before`/`after` source-observation references, and
347
- `comparisonId`/`comparisonRequestId` plus `contracts: {baselineId,
348
- contractId}` references - never an embedded `ObservationArtifact` or copied
349
- screenshot. It reuses `ClauseEvaluationResult`/`OverallVerdict`/
350
- `UnexpectedChangeResult` unchanged and contains no evaluation logic itself.
351
- `evaluationRequestId` is a deterministic function of `{baselineId,
352
- contractId, beforeObservationId, afterObservationId, comparisonRequestId}`
353
- (`frontendContractIdentity.ts#buildFrontendContractEvaluationRequestIdentity` -
354
- deliberately `comparisonRequestId`, not the fresh-per-execution
355
- `comparisonId`, so semantically identical evaluations share an identity);
356
- `evaluationId` reuses the existing generic `buildFrontendContractInstanceIdentity`
357
- unchanged. `src/artifacts/frontendContractEvaluationArtifactWriter.ts`/
358
- `frontendContractEvaluationArtifactReader.ts` persist/read it with the same
359
- atomic-write discipline as above.
360
-
361
- **Application seam**: `src/application/frontendContractEvaluationService.ts#evaluateAndPersist`
362
- calls the existing pure `evaluateFrontendContract` exactly once and - only
363
- for a structurally constructible result, whether the verdict is `PASS` or
364
- `FAIL` - persists exactly one evaluation artifact; an `{ok: false}` evaluator
365
- result (evidence could not be constructed into an evaluation at all) is never
366
- persisted as a fabricated artifact. `evaluateAndPersistFromArtifactRoots` is
367
- the future-CLI-facing wrapper: it reads two observations through the existing
368
- `readObservationArtifact` (never a second observation reader), the
369
- comparison and the two contracts through the readers above, then delegates
370
- to `evaluateAndPersist` exactly once.
371
-
372
- ## v0.5 public contract/evaluation commands (shipped as part of this release)
373
-
374
- Three public commands expose the persistence/evaluation contract above (see
375
- `docs/COMMANDS.md` for exact flags/output/exit behavior, not duplicated
376
- here):
377
-
378
- - `approve-baseline` → `frontendContractPersistenceService.ts#approveAndPersistBaseline`
379
- → validates a `PersistentBaselineContract` and its `sourceObservation`
380
- coherence against a supplied observation artifact → persists via
381
- `frontendContractArtifactWriter.ts`. The only baseline-approval act in the
382
- observer.
383
- - `save-change-contract` → `frontendContractPersistenceService.ts#persistPerChangeContract`
384
- → validates a `PerChangeContract` (rejecting a baseline contract, an
385
- authored `unexpected` category, or any other structural violation) →
386
- persists via the same writer. Persistence only, never approval.
387
- - `evaluate-contract` → `frontendContractEvaluationService.ts#evaluateAndPersistFromArtifactRoots`
388
- → `evaluateFrontendContract` exactly once → `frontendContractEvaluationArtifactWriter.ts`
389
- exactly once. `--enforce` affects only the process exit status for an
390
- already-persisted `FAIL` verdict.
391
-
392
- No command infers baseline approval or supersession automatically - not
393
- `compare`, not a `PASS` evaluation, not any artifact writer.
394
-
395
- This full command sequence is proven against real Chromium observations (not
396
- hand-constructed artifacts) - see "v0.5 real-browser workflow proof" below.
397
-
398
- ## v0.5 real-browser workflow proof (shipped as part of this release)
399
-
400
- `tests/browser/cliFrontendContracts.test.ts` and
401
- `scripts/dev/builtCliFrontendContractsBrowserSmoke.mjs` drive the complete
402
- `observe` → `approve-baseline` → `save-change-contract` → `observe` →
403
- `compare` → `evaluate-contract` sequence against a real disposable local HTTP
404
- fixture and real Chromium, proving two scenarios:
405
-
406
- - a fully successful contract change - a real observed navigation-width
407
- decrease and workspace-width increase, both satisfying their authored
408
- `requested`/`expected-dependent` clauses, an unchanged `protected` rail
409
- width, and an unclipped `preserved` navigation - overall `PASS`;
410
- - the "milestone signature" failure - the same locally successful requested
411
- change (navigation shrinks, workspace expands, both still `pass`)
412
- co-occurring with a genuine `protected` right-rail width regression (a real
413
- `resized` comparison difference) and a genuine `preserved` navigation
414
- clipping regression (a real `clipping-changed` difference, `not-clipped` →
415
- `clipped`) - overall `FAIL`.
416
-
417
- Both scenarios confirm: `--enforce` changes only the process exit status
418
- (`0` without it, nonzero with it) for the identical persisted
419
- `evaluationRequestId`/`clauseResults`; every source observation and
420
- comparison artifact is byte-identical before and after evaluation; the
421
- evaluation directory contains `manifest.json` only (no copied screenshot);
422
- and no operational filesystem path is ever serialized into a persisted
423
- manifest. This is real-browser evidence layered on top of the CLI-level
424
- proof in `tests/unit/cliFrontendContracts.test.ts` and the Chromium-free
425
- `scripts/dev/builtCliFrontendContractsSmoke.mjs` - it does not replace them.
426
-
427
- ## v0.6 bounded agent context and correlation contract (released as `0.6.0`)
428
-
429
- **Current status: released as package version `0.6.0`, tag `v0.6.0`, from
430
- the canonical `canonicalization/v0.6` lineage.** Bounded-agent-context is a new, independent
431
- artifact-kind family, schema `1.0.0` (`BOUNDED_AGENT_CONTEXT_ARTIFACT_KIND =
432
- "my-frontend-observer/bounded-agent-context"`) - never a bump to
433
- observation/comparison/frontend-contract/evaluation schemas, which remain
434
- `1.2.0`/`1.0.0`/`1.0.0`/`1.0.0` respectively. Unlike those families, there is
435
- **no disk artifact writer/reader** for bounded-agent-context: it is a pure
436
- programmatic contract and derivation layer, exported from `src/index.ts`
437
- only.
438
-
439
- **Bounded runtime projection**
440
- (`src/domain/boundedAgentContextProjection.ts#projectBoundedAgentContext`)
441
- produces a `BoundedRuntimeTargetProjection` from already-captured v0.1-v0.5
442
- evidence, containing: page/viewport identity; stable target identities;
443
- important geometry and runtime behavior; layout/behavior relationships;
444
- before/after differences; contract clause results; requested/expected-
445
- dependent/protected/preserved scope - reusing `src/domain/
446
- frontendContracts.ts`'s existing clause types verbatim, never a
447
- reimplementation; diagnostics; screenshot/artifact references; provenance;
448
- and explicit `OmissionRecord`/`TruncationRecord` metadata with bounded
449
- aggregate-cap summarization once a limit is reached.
450
-
451
- **Adequacy**: every projection carries an `Adequacy` value
452
- (`adequate`/`partial`/`inadequate`) plus a structured, closed
453
- `ADEQUACY_REASON_CODES` vocabulary - evidence existing is not itself
454
- adequacy; a required omission or an `incomparable`/unavailable upstream
455
- source is reflected honestly rather than silently reported as sufficient.
456
-
457
- **Runtime/static correlation**
458
- (`src/domain/boundedAgentContextCorrelation.ts#deriveRuntimeStaticCorrelations`/
459
- `attachRuntimeStaticCorrelations`) evaluates each stable runtime target
460
- against caller-supplied candidate static-evidence records into exactly one of
461
- three outcomes: `correlated`, `ambiguous` (multiple competing candidates,
462
- all preserved and visible - never silently resolved to one), or
463
- `unavailable` (no supported candidate). The module accepts only plain,
464
- already-retrieved candidate records and has no dependency on
465
- `@dailephd/my-dev-kit` - the audit preceding implementation found no generic
466
- static-side retrieval capability actually missing (see `docs/ROADMAP.md` v0.6
467
- "Dependency direction"). A runtime target identity is carried through
468
- verbatim; correlation never produces a `sourceOwner`/`causedBy`-shaped field,
469
- so a stable runtime identity is never silently reported as source ownership.
470
-
471
- **Identity**: `src/domain/boundedAgentContextIdentity.ts#buildBoundedAgentContextRequestIdentity`/
472
- `buildBoundedAgentContextInstanceIdentity` follow the same
473
- canonicalize+sha256(+opaque-nonce) pattern as `comparisonIdentity.ts`/
474
- `frontendContractIdentity.ts`: a deterministic logical identity distinct from
475
- a fresh per-execution instance identity.
476
-
477
- **Export/public boundary**: `src/index.ts` exports the complete
478
- bounded-agent-context/correlation type and function surface as a
479
- programmatic library contract. There is no CLI command (`observe`/`compare`/
480
- `approve-baseline`/`save-change-contract`/`evaluate-contract` remain the only
481
- public commands) and no orchestrator/lab code in this repository - bounded
482
- runtime-evidence consumption by `my-dev-kit-orchestrator` and exact
483
- readers/fixtures/evaluation in `my-dev-kit-lab` are separate sibling-
484
- repository deliverables outside `my-frontend-observer`'s public surface.
485
-
486
- **Compatibility evidence**: cross-repository neutral verification (observer
487
- `514bf3bb513764815a0a5b9e508d5836aa7d7fd8`, orchestrator `9473e4c`, lab
488
- `271e72c`) passed with 6/6 requirement coverage and no known product
489
- blockers; on the canonical worktree, `npm run typecheck`, `npm run lint`,
490
- `npm test` (627 tests), `npm run test:browser` (120 tests), `npm run
491
- test:security`, `npm run build`, and `npm run check:docs` all pass.
492
-
493
- ## v0.7 external visual-reference contract direction (released as `0.7.0`; v0.8 viewer released as `0.8.0`; v0.9 released as `0.9.0`; v0.10 still future)
494
-
495
- External visual-reference support is released as package version `0.7.0`
496
- (see "v0.7 Prompt 1" through "v0.7 Prompt 8" below for the exact contract).
497
- The exact public type names, artifact kinds, schema versions, persistence
498
- layout, and command/programmatic entry points were designed during v0.7
499
- implementation from current repository precedent, following the constraints
500
- below. v0.8 (released as package version `0.8.0` - see
501
- `docs/CURRENT_STATE.md`) has preserved them. v0.9 (implemented, not yet
502
- released) preserves them too. v0.10 remains future and must continue to
503
- preserve them.
504
-
505
- **Distinct evidence domain**: an external reference is desired-design evidence,
506
- not an `ObservationArtifact` and not the "before" side of a v0.4
507
- `ComparisonArtifact`. Reference design vs candidate is distinct from both
508
- before vs after comparison and frontend-contract evaluation. The implementation
509
- must not fake this distinction by wrapping a raster image in an observation
510
- shape.
511
-
512
- **Reference identity and provenance**: a future reference contract must preserve
513
- a deterministic logical reference identity/version where appropriate, source
514
- image reference plus dimensions/format, provenance, bounded region definitions,
515
- applicable viewport/theme/application-state identity, authored design intent,
516
- relationship/style evidence where supported, limits/diagnostics, and approval/
517
- supersession history. Operational filesystem paths must not become semantic
518
- identity. A raw imported image never silently becomes an approved active
519
- reference.
520
-
521
- **Reference regions and runtime targets stay distinct**: a reference region
522
- must have its own identity and coordinate semantics. Reference-region to runtime-
523
- target association must be explicit and capable of representing ambiguity or
524
- unavailability. Runtime target identity and reference identity must never
525
- silently become static source ownership; static association still goes through
526
- the v0.6 runtime/static correlation boundary.
527
-
528
- **Applicability before fidelity**: viewport, theme, application state, and
529
- other selected compatibility dimensions must be evaluated before ordinary
530
- reference/candidate differences are produced. If the reference and candidate
531
- represent different intended states, the result must be explicitly incompatible
532
- or incomparable rather than filled with fabricated visual failures. Planning
533
- should reuse or extend the canonical v0.4 comparability conventions where they
534
- mean the same thing rather than invent an unrelated reference-only state model.
535
-
536
- **Canonical contract semantics remain authoritative**: executable reference-
537
- derived requirements must map into the existing v0.5 authored categories
538
- `requested`, `expected-dependent`, `protected`, or `preserved`. The derived-only
539
- `unexpected` classification remains derived-only. Informational or unassessed
540
- reference evidence may stay outside executable contract evaluation until
541
- explicitly promoted. A second reference-only PASS/FAIL taxonomy is forbidden.
542
-
543
- **Tolerance separation**: reference-fidelity tolerances are not automatically
544
- the same as v0.4 `ComparisonConfig.geometryTolerancePx` or v0.5 contract
545
- tolerances. Planning must define property-specific semantics for reference
546
- geometry, spacing, selected style evidence, text/font rendering differences,
547
- asset-sensitive regions, and optional image similarity. One global pixel-perfect
548
- threshold is not an acceptable contract.
549
-
550
- **Structured evidence first**: geometry, relationships, authored requirements,
551
- applicability, provenance, and selected bounded style/asset evidence remain
552
- inspectable primary evidence. Screenshot-region or image-similarity evidence may
553
- supplement them where reliable, but pixel similarity alone must not determine
554
- success and must never override active baseline/per-change contracts.
555
-
556
- **Bounded correction evidence**: future reference/candidate results must support
557
- a bounded projection suitable for coding-agent correction, such as reference
558
- measurement, candidate measurement, delta, failed relationship/style condition,
559
- relevant reference/runtime identities, provenance, and active protected/
560
- preserved constraints. Heavy reference image bytes should be referenced, not
561
- copied into every downstream context packet.
562
-
563
- **Approval and supersession**: reference import, reference approval, baseline
564
- approval, reference supersession, and baseline supersession are separate acts.
565
- A reference-fidelity `PASS`, a frontend-contract `PASS`, or a successful
566
- before/after comparison must not silently approve or replace any reference or
567
- baseline.
568
-
569
- The v0.8 viewer, released as package version `0.8.0`, consumes this v0.7
570
- reference/evaluation contract exactly as required - it creates no UI-only
571
- reference model (see `docs/ARCHITECTURE.md` "v0.8 Batch 5"/"v0.8 Batch 6"
572
- and `docs/reports/v0.8-reference-candidate-inspection-batch5.md`). v0.9
573
- annotations (released in `0.9.0`) originate from runtime
574
- screenshots or external references, preserve which source identity and
575
- coordinate system they belong to, and feed the same canonical contract and
576
- reference semantics - see "v0.9 visual annotation contract" below. v0.10
577
- combines both entry modes into the full correction/approval workflow.
578
-
579
- ## v0.9 visual annotation contract (released in 0.9.0)
580
-
581
- v0.9 is released as `@dailephd/my-frontend-observer@0.9.0`.
582
-
583
- **Artifact**: `VisualAnnotationArtifact`, artifact kind
584
- `my-frontend-observer/visual-annotation`, schema version `1.0.0`. It stores one
585
- exact canonical source (a runtime observation, or an imported or approved
586
- external reference), the source coordinate space (runtime CSS pixels or
587
- reference-image pixels), and bounded structured items. Each item has a stable
588
- `annotationItemId`, one mark (point, rectangle, line, arrow, or note), an
589
- optional explicit association, and an interpretation. A revision sets
590
- `supersedesAnnotationId` and never rewrites its parent. The overlay SVG is
591
- derived from the artifact and verified before it is served.
592
-
593
- **An annotation is not a contract.** Saving an annotation never creates a
594
- contract clause, a reference requirement, or a PASS/FAIL rule. Marks and
595
- visible pixels are evidence. Overlap between a mark and a target or region is
596
- not ownership and never creates an association.
597
-
598
- **Interpretation states**: `uninterpreted`, `candidate`, and `confirmed`.
599
- Confirmation is explicit and records `confirmedAt`. Editing the mark,
600
- association, or intent of a confirmed item withdraws the confirmation.
601
-
602
- **Runtime intent to contract**:
603
-
604
- - Only selected, confirmed, supported runtime intent is promoted. Promotion
605
- creates one normal canonical `PerChangeContract` through the existing
606
- contract persistence service.
607
- - Supported mappings use the existing `ContractPrimitive` vocabulary only.
608
- `move` maps to `property-increases`/`property-decreases` on `x` or `y`.
609
- `resize` maps to `property-increases`/`property-decreases` on `width` or
610
- `height`. `preserve` maps to `property-unchanged-within-tolerance` for a
611
- target property, or to `relationship-unchanged` for an explicitly associated
612
- canonical relationship.
613
- - Categories are the canonical `requested`, `expected-dependent` (with a
614
- required `required` or `permitted` mode), `protected`, and `preserved`.
615
- `unexpected` is never authored; it stays evaluator-derived.
616
- - `remove` can be confirmed and saved, but it is not promotable in v0.9. The
617
- contract vocabulary has no target-absent primitive, and no approximate
618
- clause is fabricated.
619
- - `inspect` intent and notes are informational and never promoted.
620
- - Each clause's `supportingEvidence` records the annotation source and item
621
- paths. Promotion never activates the contract unless explicitly requested,
622
- and it never approves a baseline.
623
-
624
- **Reference intent to a new reference revision**:
625
-
626
- - Only selected, confirmed `reference-region` (`create` or `refine`) and
627
- `reference-requirement` items are materialized. `inspect` and
628
- `asset-sensitive` intent stay informational.
629
- - The result is a new imported `ExternalReferenceArtifact`, created through the
630
- existing canonical import service. It supersedes the selected source
631
- reference (the approved reference id when the source is approved) and has
632
- lifecycle `imported`. It is never approved automatically.
633
- - Source regions keep their order. A refine replaces the rectangle of an
634
- existing source region and keeps its id. Creates are appended in selection
635
- order. Duplicate creates, duplicate refines, and create-then-refine in one
636
- request are rejected.
637
- - Source requirements stay first with identical recomputed ids. Selected
638
- requirements are appended in selection order and validated against the final
639
- regions. A selected relationship requirement must still hold for the final
640
- geometry. Measurement requirements store only the subject and tolerance.
641
- - Applicability, label, and the exact source image bytes are preserved. The
642
- source reference is never modified, and project reference acceptance is not
643
- changed.
644
-
645
- Existing contract, reference relationship, adequacy, and fidelity evaluators
646
- remain the only source of verdicts. v0.9 adds no annotation evaluator.
647
-
648
- ## Approved v0.1 design inputs
649
-
650
- The historical greenfield scaffold plan recorded these v0.1 design decisions:
651
-
652
- - artifact kind `my-frontend-observer/observation`;
653
- - schema version `1.0.0`, independent of package version;
654
- - one portable directory containing `manifest.json`, `evidence.json`, and
655
- `screenshot.png`;
656
- - evidence states `available`, `unavailable`, `not-applicable`, and `partial`;
657
- - evidence sources `browser`, `computed-browser`, and `derived`;
658
- - bounded explicitly requested targets, provenance, diagnostics, completion
659
- state, limits, and relative artifact references.
660
-
661
- These were planning inputs only at the time they were recorded. As shown in
662
- "Current contracts" above, the implemented contract matches them except for
663
- the file layout: there is no separate `evidence.json` - page/target evidence
664
- is embedded directly inside `manifest.json`.
665
-
666
- Comparison and relationship contracts belong to v0.4, and canonical
667
- change-scope contracts belong to v0.5 - see "v0.5 frontend contract and
668
- evaluation" above for the full shipped contract model, identity, evaluation
669
- engine, persistence, baseline approval, and CLI exposure. Bounded
670
- agent-context and runtime/static correlation contracts are v0.6 - see "v0.6
671
- bounded agent context and correlation contract" above for the full released
672
- model. The text/config-driven coding-agent review plus non-graphical external
673
- visual-reference foundation is v0.7 - see "v0.7 Prompt 1 external-reference
674
- artifact contract" below for the foundation layer implemented so far. Viewer
675
- consumption of that reference model is released in v0.8 (see "v0.7 external
676
- visual-reference contract direction" above); dual-context annotation follows
677
- in v0.9; both visual entry modes converge with the existing workflow in
678
- v0.10.
679
-
680
- ## v0.7 Prompt 1 external-reference artifact contract
681
-
682
- Released as `0.7.0`. This is the foundation layer only: identity,
683
- provenance, bounded image metadata, and a two-state lifecycle for one
684
- externally supplied design-reference image. It implements no region,
685
- geometry, relationship, requirement, tolerance, binding, or fidelity-
686
- evaluation contract - those belong to later v0.7 prompts.
687
-
688
- An external reference is a distinct evidence root, not a variant of
689
- `ObservationArtifact`: it never reuses `ARTIFACT_KIND`/`SCHEMA_VERSION`
690
- (observation), `COMPARISON_ARTIFACT_KIND`, or `CONTRACT_ARTIFACT_KIND`, and
691
- those existing types gain no new field from this contract.
692
-
693
- ```ts
694
- const EXTERNAL_REFERENCE_ARTIFACT_KIND = 'my-frontend-observer/external-reference';
695
- const EXTERNAL_REFERENCE_SCHEMA_VERSION = '1.0.0'; // independent of package.json version and every other family's schema version
696
-
697
- type ExternalReferenceImageFormat = 'png' | 'jpeg' | 'webp';
698
-
699
- interface ExternalReferenceImageReference {
700
- path: string; // bare relative filename within the artifact's own directory
701
- format: ExternalReferenceImageFormat;
702
- width: number;
703
- height: number;
704
- byteLength: number;
705
- sha256: string; // identity-bearing content hash of the raw image bytes
706
- }
707
-
708
- // Points back to the imported artifact that owns the image, without copying its bytes - mirrors ComparisonSourceObservationReference.
709
- interface ExternalReferenceSourceReference {
710
- referenceId: string;
711
- referenceRequestId: string;
712
- producer: { name: 'my-frontend-observer'; version: string };
713
- schemaVersion: '1.0.0';
714
- image: ExternalReferenceImageReference;
715
- }
716
-
717
- // Exactly two persisted states - no literal 'superseded' variant (see below).
718
- type ExternalReferenceLifecycleState = { state: 'imported' } | { state: 'approved'; approvedAt: string };
719
-
720
- interface ExternalReferenceArtifactBase {
721
- artifactKind: 'my-frontend-observer/external-reference';
722
- schemaVersion: '1.0.0';
723
- referenceRequestId: string; // deterministic logical identity - shared by an imported artifact and every artifact produced by approving it
724
- referenceId: string; // fresh per-persisted-instance identity
725
- producer: { name: 'my-frontend-observer'; version: string };
726
- provenance: { importedAt: string; label?: string };
727
- supersedesReferenceId?: string; // explicit, forward-only supersession of a prior reference's referenceId
728
- diagnostics: Diagnostic[];
729
- completion: CompletionState;
730
- }
731
-
732
- // lifecycle.state === 'imported': owns the image.
733
- interface ImportedExternalReferenceArtifact extends ExternalReferenceArtifactBase {
734
- lifecycle: { state: 'imported' };
735
- image: ExternalReferenceImageReference;
736
- }
737
-
738
- // lifecycle.state === 'approved': references, never copies, the imported artifact's image.
739
- interface ApprovedExternalReferenceArtifact extends ExternalReferenceArtifactBase {
740
- lifecycle: { state: 'approved'; approvedAt: string };
741
- sourceReference: ExternalReferenceSourceReference;
742
- }
743
-
744
- type ExternalReferenceArtifact = ImportedExternalReferenceArtifact | ApprovedExternalReferenceArtifact;
745
- ```
746
-
747
- Key rules:
748
-
749
- - `referenceRequestId` is a pure function of `{imageSha256, format, width,
750
- height, supersedesReferenceId}` only - never a filesystem path, output
751
- location, label, or timestamp. Byte-identical image content imported from a
752
- different operational root produces the same `referenceRequestId`;
753
- changing any of those fields changes it.
754
- - `referenceId` is fresh (nonce-based) on every persisted write, including
755
- every approval of an already-imported reference.
756
- - Importing an image never approves it (`lifecycle.state` is always
757
- `'imported'` immediately after import, regardless of a supplied label or
758
- supersession target). Approval is a single explicit act
759
- (`approveExternalReference`, mirroring `approveAndPersistBaseline`) that
760
- refuses anything not currently in the `'imported'` state.
761
- - Approving persists a *new* artifact instance (same `referenceRequestId`,
762
- fresh `referenceId`) carrying a `sourceReference` back to the imported
763
- artifact - it never mutates the imported artifact's own manifest, and never
764
- copies the image bytes a second time.
765
- - Supersession is represented only as a forward pointer
766
- (`supersedesReferenceId` on the newer artifact); there is deliberately no
767
- literal `'superseded'` lifecycle state, so an existing persisted artifact's
768
- own manifest is never rewritten - immutability holds unconditionally rather
769
- than depending on careful mutation discipline.
770
- - Supported formats are frozen to exactly `png`/`jpeg`/`webp`, detected from
771
- header/magic bytes only (never a caller-declared file extension), bounded
772
- to `EXTERNAL_REFERENCE_MAX_IMAGE_BYTES` (20,000,000 bytes) and
773
- `[EXTERNAL_REFERENCE_MIN_DIMENSION_PX, EXTERNAL_REFERENCE_MAX_DIMENSION_PX]`
774
- (`[1, 8192]`) pixels per side. No OCR, no raster decode, no computer
775
- vision, no automatic region detection.
776
-
777
- Persisted as `<outputLocation>/<referenceId>/manifest.json` (+
778
- `reference.<ext>` for an `'imported'` artifact only), via the same atomic
779
- temp-dir-then-rename discipline as every other artifact family
780
- (`src/artifacts/externalReferenceArtifactWriter.ts` /
781
- `externalReferenceArtifactReader.ts`). CLI: `import-reference <image-file>
782
- --output <dir> [--label] [--supersedes <root>]` and `approve-reference
783
- --reference <root> --output <dir> [--supersedes <root>]`.
784
-
785
- ## v0.7 Prompt 2 explicit reference regions and relationships
786
-
787
- Released as `0.7.0`. Additive extension of the Prompt 1 contract above:
788
- one new, optional `regions?: ReferenceRegion[]` field on
789
- `ExternalReferenceArtifact` (both lifecycle variants), plus a pure,
790
- non-persisted relationship-derivation capability. No schema version bump -
791
- `EXTERNAL_REFERENCE_SCHEMA_VERSION` remains `'1.0.0'`, because the field is
792
- genuinely optional/additive and every Prompt 1 artifact (which predates this
793
- field entirely) remains valid without it.
794
-
795
- ```ts
796
- // domain/externalReferenceRegions.ts
797
- interface ReferenceRegionRectangle { x: number; y: number; width: number; height: number; }
798
- interface ReferenceRegion { id: string; rectangle: ReferenceRegionRectangle; }
799
-
800
- // Pure derived geometry - never persisted, always recomputed, so it can never drift from the rectangle above.
801
- interface ReferenceRegionGeometry {
802
- x: number; y: number; width: number; height: number;
803
- right: number; bottom: number; centerX: number; centerY: number;
804
- }
805
-
806
- const REFERENCE_REGION_ID_PATTERN = /^[A-Za-z0-9_-]{1,64}$/; // same convention as request/request.ts's target-name pattern
807
- const MAX_REFERENCE_REGIONS = 20; // same bound value as request/request.ts's MAX_TARGETS - independently owned, coincidentally equal
808
- ```
809
-
810
- Region coordinate semantics: origin at the reference image's top-left
811
- corner, x increasing rightward, y increasing downward, unit is
812
- reference-image pixels (explicitly not CSS pixels - a static image has no
813
- CSS box model), coordinates may be fractional. A region's rectangle must lie
814
- entirely within its owning image's own already-validated
815
- width/height - out-of-bounds geometry is rejected outright, never clamped.
816
-
817
- Key rules:
818
-
819
- - Only `{x, y, width, height}` is canonical/authored. `right`, `bottom`,
820
- `centerX`, `centerY` are pure calculations over it
821
- (`deriveReferenceRegionGeometry`) - never a second, potentially-drifting
822
- stored copy of the same fact.
823
- - Region content is identity-bearing:
824
- `buildExternalReferenceRequestIdentity` gained an additive, optional
825
- trailing `regions` parameter. Omitting it entirely (every Prompt 1 call
826
- site, and any Prompt 2 call that legitimately has no regions) produces the
827
- byte-identical hash Prompt 1 already produced - the parameter is left out
828
- of the hashed view rather than defaulted to `null`, unlike
829
- `supersedesReferenceId`. Authored region order participates in identity
830
- (arrays are never reordered by the shared `canonicalize()`), mirroring
831
- `domain/identity.ts`'s treatment of configured targets.
832
- - Region IDs are unique case-insensitively within one artifact (mirroring
833
- `request/request.ts`'s target-name dedup convention exactly).
834
- - `import-reference` gained an optional `--regions-file <json-file>` of the
835
- form `{ "regions": [...] }` (same object-root-wrapper convention as
836
- `--targets-file`); a legacy invocation without it behaves exactly as in
837
- Prompt 1. `approve-reference` carries an imported artifact's `regions`
838
- forward verbatim (never re-validated, never re-derived, never dropped) -
839
- approval never adds, removes, or edits regions.
840
- - One new diagnostic code, `invalid-reference-region` (error), covers every
841
- region-validation failure (missing/duplicate/malformed id,
842
- non-finite/negative/zero geometry, out-of-image-bounds, over the bounded
843
- region count) - deliberately not split into several codes, per the
844
- "don't proliferate diagnostics" convention.
845
-
846
- Reference-region relationships (`domain/externalReferenceRegionRelationships.ts`)
847
- reuse the exact same pure, tolerance-aware geometry predicates that
848
- `domain/relationships.ts#deriveLayoutRelationships` uses for runtime targets
849
- (`horizontalOrderOf`/`verticalOrderOf`/`areaOverlapOf`/`relativeWidthOf`/
850
- `geometricFitOf`/`verticalSequenceOf`, now exported additively from that
851
- module with unchanged formulas) and the same `PairwiseRelationshipKind`
852
- vocabulary and `EvidenceReference` type - never a duplicated or
853
- reinterpreted copy. Only the six geometry-only families apply (horizontal
854
- order, vertical order, area overlap, relative width, geometric fit, vertical
855
- sequencing); DOM containment, scroll ownership, runtime visibility, and
856
- page-width-vs-viewport are runtime/browser concepts with no reference-image
857
- equivalent and are not reused. `fits-inside`/`does-not-fit-inside` is
858
- geometry-only fit - it never claims DOM containment, which an external image
859
- cannot expose.
860
-
861
- ```ts
862
- interface ReferenceRegionRelationship {
863
- kind: PairwiseRelationshipKind;
864
- subjectRegion: string; // deliberately distinct field name from PairwiseLayoutRelationship's subjectTarget
865
- relatedRegion: string;
866
- evidence: EvidenceReference[]; // e.g. { path: 'regions.header.rectangle' } - never a targetEvidence/browser path
867
- }
868
- ```
869
-
870
- Relationships are **not persisted** on the artifact - `deriveReferenceRegionRelationships(referenceRequestId, regions, options)`
871
- is a pure, deterministic, synchronous function any caller (a future prompt,
872
- a test) calls on demand against an artifact's own `regions` field, avoiding
873
- any possibility of a persisted relationship graph drifting from the region
874
- data it was derived from. Bounded at `MAX_REFERENCE_REGIONS` regions ->
875
- `MAX_REFERENCE_REGION_PAIRS` pairs `x` 6 families =
876
- `MAX_REFERENCE_REGION_RELATIONSHIP_RECORDS` records maximum - the same
877
- bounding shape as `relationships.ts`'s `MAX_PAIRWISE_RELATIONSHIP_RECORDS`.
878
- This is a maximum capacity, never a required minimum region count - there is
879
- no contract requiring any specific number of authored regions.
880
-
881
- A reference relationship is a fact about the reference image's geometry
882
- only. It is not a design requirement, not a pass/fail verdict, and does not
883
- claim a runtime target or source owner exists - see
884
- `docs/WORKFLOWS.md` "Current external-reference foundation workflow" for
885
- where those later concepts (Prompt 3+) will attach.
886
-
887
- ## v0.7 Prompt 3 selected design requirements, tolerance semantics, and reference-evidence adequacy
888
-
889
- Released as `0.7.0`. Additive extension of the Prompt 1/2 contracts
890
- above: one new, optional `requirements?: ExternalReferenceRequirement[]`
891
- field on `ExternalReferenceArtifact` (both lifecycle variants). No schema
892
- version bump - same reasoning as Prompt 2's `regions` field.
893
-
894
- **Central distinction**: a region's geometry is REFERENCE EVIDENCE -
895
- everything visibly/measurably present in the image. A requirement is
896
- SELECTED DESIGN INTENT - only what the user/configuration explicitly chose
897
- as mattering for later candidate evaluation. Nothing in this repository ever
898
- turns a region property or a derived relationship into a requirement
899
- automatically.
900
-
901
- ```ts
902
- // domain/externalReferenceRequirements.ts
903
-
904
- // Reused directly from domain/frontendContracts.ts - not reinvented as a
905
- // "reference-only" taxonomy; that type carries no runtime-only coupling.
906
- // 'unexpected' remains impossible to author (not a member of this union).
907
- type AuthoredChangeScopeCategory = 'requested' | 'expected-dependent' | 'protected' | 'preserved';
908
- type ExpectedDependentMode = 'required' | 'permitted'; // required only (and exactly) when category === 'expected-dependent'
909
-
910
- type ReferenceRequirementRegionProperty = 'x' | 'y' | 'width' | 'height' | 'right' | 'bottom' | 'centerX' | 'centerY'; // exactly ReferenceRegionGeometry's own fields
911
- type ReferenceRequirementMeasurement = 'vertical-gap' | 'horizontal-gap' | 'center-x-delta' | 'center-y-delta' | 'left-edge-delta' | 'right-edge-delta';
912
-
913
- type ReferenceRequirementSubject =
914
- | { kind: 'region-property'; region: string; property: ReferenceRequirementRegionProperty }
915
- | { kind: 'region-relationship'; subjectRegion: string; relatedRegion: string; relationship: PairwiseRelationshipKind } // reused from relationships.ts - geometry-only families only
916
- | { kind: 'region-measurement'; subjectRegion: string; relatedRegion: string; measurement: ReferenceRequirementMeasurement };
917
-
918
- // Deliberately NOT a reuse of frontendContracts.ts's ContractTolerance: that
919
- // type's 'absolute-px' is implicitly runtime/CSS pixels. Reference-image
920
- // pixels are a distinct, explicitly-labeled unit - nothing here assumes
921
- // 1 reference pixel = 1 CSS pixel (Prompt 6 will need an explicit mapping).
922
- type ReferenceRequirementTolerance = { kind: 'exact' } | { kind: 'absolute-reference-px'; amount: number } | { kind: 'percent'; amount: number };
923
-
924
- interface ExternalReferenceRequirement {
925
- requirementId: string; // system-computed from {subject, category, expectedDependentMode, tolerance} only - never authored
926
- category: AuthoredChangeScopeCategory;
927
- expectedDependentMode?: ExpectedDependentMode;
928
- subject: ReferenceRequirementSubject;
929
- tolerance?: ReferenceRequirementTolerance; // required for region-property/region-measurement; must be absent for region-relationship
930
- }
931
- ```
932
-
933
- Key rules:
934
-
935
- - Requirement identity (`requirementId`) is always system-computed
936
- (`buildReferenceRequirementIdentity`, mirroring
937
- `frontendContractIdentity.ts#buildClauseIdentity`'s exact shape) - the raw
938
- authored input (`RawReferenceRequirement`) has no `requirementId` field at
939
- all, and supplying one is a validation error. Unlike v0.5's
940
- `BaselineClause`/`PerChangeClause` (which need an author-visible `clauseId`
941
- for cross-document `supersedesBaselineClauseIds` references), Prompt 3
942
- requirements have no cross-document reference need yet, so trusting an
943
- authored id would only invite drift between a user-typed id and the
944
- content it claims to identify.
945
- - The reference-side expected value/relationship is never stored on the
946
- requirement or the artifact - `deriveReferenceRequirementExpectation()` is
947
- a pure function computed on demand from the artifact's own `regions`,
948
- eliminating the exact drift risk of persisting e.g. `width: 424` alongside
949
- a region whose rectangle could (in principle) later disagree with it.
950
- - A requirement referencing a region id that does not exist in the
951
- artifact's own `regions` is a **structural validation failure** (rejected
952
- at construction/import time), never merely "unavailable" reference
953
- evidence - `isValidReferenceRequirements` checks this before any
954
- requirement reaches adequacy computation.
955
- - **Duplicate/conflicting subject rule**: no two requirements in one
956
- collection may share the same structural subject (same region+property,
957
- or the same unordered region pair + relationship, or + measurement),
958
- regardless of category. This single rule covers both "duplicate
959
- requirement" and "conflicting categories on the same subject" (e.g. the
960
- same region/property authored as both `requested` and `protected`) -
961
- v0.5's `evaluateFrontendContract#primitivesConflict` is a *runtime-
962
- evaluation-time* detector (it needs before/after `ObservationArtifact`
963
- evidence that does not exist yet at this stage) and could not be reused
964
- safely; Prompt 3 restricts invalid combinations at authoring time instead,
965
- per the documented precedent-review outcome.
966
- - Bounded at `MAX_REFERENCE_REQUIREMENTS` (50) requirements per artifact -
967
- a maximum capacity, never a required minimum (there is no contract
968
- requiring any specific number of authored requirements).
969
- - One new diagnostic code, `invalid-reference-requirement` (error), covers
970
- every requirement-authoring validation failure - deliberately not split
971
- further, per the "don't proliferate diagnostics" convention already used
972
- for `invalid-reference-region`.
973
- - `import-reference` gained an optional `--requirements-file <json-file>`
974
- (`{ "requirements": [...] }`, same object-root-wrapper convention as
975
- `--regions-file`/`--targets-file`); `approve-reference` carries an
976
- imported artifact's `requirements` forward verbatim (never re-validated,
977
- never re-derived, never dropped), exactly mirroring how it already
978
- handles `regions`.
979
-
980
- **Reference-evidence adequacy** (`deriveReferenceRequirementAdequacy(regions, requirements)`)
981
- answers only "does the reference definition itself contain enough evidence
982
- to understand every selected requirement?" - never "does a runtime
983
- target/candidate exist" (that is Prompt 4/5's responsibility). It is its own
984
- small, reference-owned vocabulary (`REFERENCE_REQUIREMENT_ADEQUACY_STATES` =
985
- `'adequate' | 'partial' | 'inadequate'`, and exactly two reason codes,
986
- `no-selected-requirements` and `missing-reference-relationship-evidence`) -
987
- deliberately **not** a reuse of
988
- `boundedAgentContext.ts`'s `Adequacy`/`ADEQUACY_REASON_CODES`, which
989
- describe runtime-target/static-correlation concerns that do not exist at
990
- this stage; mislabeling reference adequacy as bounded-agent-context adequacy
991
- would conflate two genuinely different evidence domains. Zero selected
992
- requirements is explicitly `inadequate` (a region-rich, fully-valid
993
- reference is still not usable for a correction task until the user has
994
- actually selected what matters) - this is a documented product decision,
995
- not an oversight. The result is never a numeric score, always structured
996
- and inspectable, with reasons ordered deterministically by authored
997
- requirement position.
998
-
999
- ```ts
1000
- interface ReferenceRequirementAdequacy {
1001
- status: 'adequate' | 'partial' | 'inadequate';
1002
- totalRequirements: number;
1003
- evaluableRequirements: number;
1004
- unavailableRequirements: number;
1005
- reasons: { code: 'no-selected-requirements' | 'missing-reference-relationship-evidence'; requirementId?: string; detail?: string }[];
1006
- }
1007
- ```
1008
-
1009
- ## v0.7 Prompt 4 reference applicability and candidate-state compatibility
1010
-
1011
- Released as `0.7.0`. Additive extension of the Prompt 1/2/3 contracts
1012
- above: one new, optional `applicability?: ExternalReferenceApplicability`
1013
- field on `ExternalReferenceArtifact` (both lifecycle variants), one new,
1014
- optional `explicitState?: ExplicitStateDimensions` field on
1015
- `ObservationArtifact.requestConfig`, and one new pure module,
1016
- `domain/externalReferenceCompatibility.ts`, that answers a single question:
1017
- "does this external reference describe the same frontend state as this
1018
- candidate `ObservationArtifact`?" No schema version bump on either artifact
1019
- - same reasoning as Prompt 2/3's additive fields.
1020
-
1021
- **Central distinction**: this is page/state-level compatibility only -
1022
- never geometry, never fidelity, never a visual/pixel comparison, and never
1023
- region-to-runtime-target binding (Prompt 5). It answers "should a
1024
- reference-vs-candidate geometry comparison even be attempted", not "does the
1025
- candidate match the reference". Reference-evidence adequacy (Prompt 3) and
1026
- reference/candidate compatibility (Prompt 4) are deliberately independent:
1027
- a reference can be `adequate` (enough selected requirements to evaluate)
1028
- while simultaneously `incomparable` against a given candidate (wrong
1029
- viewport/theme/state), and vice versa - neither result constrains the
1030
- other.
1031
-
1032
- **State identity is always explicit, never inferred.** `theme`,
1033
- `applicationState`, and `authenticatedState` are caller/configuration-
1034
- supplied labels only. The observer never reads screenshot pixels, CSS, DOM
1035
- classes/text, URLs, source code, filenames, accessibility labels,
1036
- localStorage, or cookies to determine state - there is no automatic state
1037
- detection anywhere in this codebase, and Prompt 4 does not add any. Labels
1038
- are bounded opaque identities (`^[A-Za-z0-9_-]{1,64}$`, the same pattern
1039
- already used for target names and region ids) compared by exact,
1040
- case-sensitive string equality only - `"dark"` and `"one-dark"` are
1041
- unrelated labels, never fuzzy-matched or normalized.
1042
-
1043
- ```ts
1044
- // domain/explicitState.ts - shared by both ObservationArtifact and ExternalReferenceArtifact
1045
- type AuthenticatedState = 'authenticated' | 'unauthenticated'; // closed vocabulary - never a place for credentials/tokens/cookies/session ids
1046
- interface ExplicitStateDimensions {
1047
- theme?: string;
1048
- applicationState?: string;
1049
- authenticatedState?: AuthenticatedState;
1050
- }
1051
- // isValidExplicitStateDimensions requires at least one dimension declared and rejects any unsupported field -
1052
- // this is a bounded, closed shape, never an arbitrary Record<string, unknown> metadata bag.
1053
-
1054
- // domain/externalReferenceApplicability.ts
1055
- interface ApplicableViewport { width: number; height: number } // CSS pixels, bounds [200, 3840] mirroring request.ts's own viewport bounds
1056
- interface ExternalReferenceApplicability extends ExplicitStateDimensions {
1057
- viewport?: ApplicableViewport;
1058
- }
1059
- ```
1060
-
1061
- **Reference image size is never the same concept as applicable viewport.**
1062
- `ExternalReferenceImageReference.width/height` (Prompt 1) describes the
1063
- reference image's own pixel dimensions - a property of the image file,
1064
- detected from its header bytes. `applicability.viewport` describes the
1065
- CSS-pixel runtime viewport the design *represents* - a reference image may
1066
- be captured at any resolution or device-pixel-ratio (e.g. a 1920x1080
1067
- screenshot representing a 960x540 CSS-pixel layout at 2x DPR). Nothing in
1068
- `externalReferenceApplicability.ts` reads or derives a viewport from image
1069
- dimensions; `isValidExternalReferenceApplicability` is its own validator
1070
- (not a reuse of `isValidExplicitStateDimensions`, whose "at least one
1071
- dimension" rule would incorrectly reject a viewport-only applicability
1072
- object).
1073
-
1074
- **v0.4 comparability is reused, not duplicated.** `domain/comparison.ts`
1075
- gained four additive reason codes (`viewport-unassessed`, `theme-mismatch`,
1076
- `authenticated-state-mismatch`, `application-state-mismatch` - the
1077
- `*-unassessed` codes for theme/authenticated-state/application-state
1078
- already existed from v0.4) and two optional fields on `ComparabilityReason`
1079
- (`referenceValue?: string`, `candidateValue?: string`, populated only for a
1080
- mismatch reason). `domain/comparisonEngine.ts` gained one new exported pure
1081
- helper, `assessOptionalComparabilityDimension(mismatchCode, unassessedCode,
1082
- beforeValue, afterValue, mismatchMessage, unassessedMessage)`, extracted
1083
- from - and now used by - both v0.4's own `evaluateComparability`
1084
- (Observation-vs-Observation) and the new
1085
- `evaluateReferenceCandidateCompatibility` (Reference-vs-Observation). The
1086
- rule is identical either way: both values defined and equal -> no reason;
1087
- both defined and different -> a `blocking` mismatch reason (with
1088
- `referenceValue`/`candidateValue` populated); either value undefined ->
1089
- an `unassessed` reason. This is a genuine, additive improvement to v0.4's
1090
- own behavior: `evaluateComparability` now assesses theme/authenticated-
1091
- state/application-state as matching or blocking-mismatched whenever *both*
1092
- observations declare `requestConfig.explicitState`, rather than always
1093
- reporting them unassessed - but every historical observation pair (and any
1094
- pair where either side omits `explicitState`) retains the exact old
1095
- unassessed-only behavior, verified by the frozen `evaluateComparability`
1096
- regression test that predates this batch.
1097
-
1098
- ```ts
1099
- // domain/externalReferenceCompatibility.ts
1100
- interface ReferenceCandidateCompatibilityResult {
1101
- referenceId: string;
1102
- referenceRequestId: string;
1103
- candidateObservationId: string;
1104
- candidateRequestId: string;
1105
- compatibility: ComparabilityResult; // v0.4's own reused result type - state/reasons, never a boolean or a visual score
1106
- }
1107
- function evaluateReferenceCandidateCompatibility(reference: ExternalReferenceArtifact, candidate: ObservationArtifact): ReferenceCandidateCompatibilityResult;
1108
- ```
1109
-
1110
- Key rules:
1111
-
1112
- - Pure and synchronous - no browser, no filesystem, no network, no target
1113
- binding. Only `reference.applicability` and
1114
- `candidate.requestConfig.viewport`/`candidate.requestConfig.explicitState`
1115
- are consulted; reference regions/requirements are never read here (a
1116
- distinct, separate concern - see Prompt 3 above).
1117
- - A dimension the reference constrains but the candidate entirely omits
1118
- (or vice versa) is `unassessed`, never treated as compatible-by-default
1119
- and never fabricated as a mismatch - fail-closed, honest non-assessment.
1120
- - A reference that declares no `applicability` at all produces a fully
1121
- `unassessed` (never automatically `incomparable`, never automatically
1122
- `comparable` beyond "no blocking reasons found") result across all four
1123
- dimensions - Prompt 1/2/3 references remain fully usable, just
1124
- unassessed for compatibility until applicability is authored.
1125
- - No automatic persisted compatibility artifact. This is a pure
1126
- programmatic result, produced on demand by an application/CLI caller
1127
- that already holds both a reference and a candidate artifact - inventing
1128
- a new persisted artifact kind for a value this cheap to recompute would
1129
- add drift risk (a candidate/reference re-imported later could silently
1130
- disagree with a stale persisted compatibility record) with no
1131
- corresponding benefit; this may be revisited only if a later prompt's
1132
- architecture proves persistence necessary.
1133
- - Identity impact: `buildExternalReferenceRequestIdentity` gained a final
1134
- optional `applicability` parameter (omitted, never `null`, when absent -
1135
- byte-identical to Prompt 1/2/3 hashes for every call that doesn't supply
1136
- it); `buildRequestIdentity` gained a final optional `explicitState`
1137
- parameter with the identical omission convention. Neither identity
1138
- function ever takes a file path.
1139
- - CLI: `import-reference` gained an optional `--applicability-file
1140
- <json-file>` (the raw, unwrapped applicability object - not a
1141
- `{ "requirements": [...] }`-style wrapper, since applicability is a
1142
- single object rather than a named list); `observe` gained an optional
1143
- `--state-file <json-file>` (the raw, unwrapped `ExplicitStateDimensions`
1144
- object). Both follow the existing `--scroll-scenario-file` convention
1145
- exactly: relative paths resolve from the current working directory, the
1146
- path itself is never persisted or included in any identity, and CLI code
1147
- owns only flag syntax/file reading/JSON parsing/object-root validation -
1148
- all semantic validation happens in the domain layer.
1149
-
1150
- ## v0.7 Prompt 5 explicit reference-region <-> runtime-target binding
1151
-
1152
- Released as `0.7.0`. One new pure domain module,
1153
- `domain/externalReferenceRuntimeBinding.ts`, answering "which stable
1154
- observer runtime target, if any, does this candidate observation resolve
1155
- for each explicitly declared reference region?" No new field is added to
1156
- either `ExternalReferenceArtifact` or `ObservationArtifact` - both remain
1157
- exactly as Prompt 4 left them - and no schema version bump on either.
1158
-
1159
- **Two identity domains, kept strictly separate.** A binding declaration
1160
- names a Prompt 2 `ReferenceRegion.id` and a v0.2 `NamedTarget.name` (the
1161
- stable observer runtime target identity established since v0.2 - never a
1162
- CSS selector, DOM node handle, source file, React component name, or
1163
- my-dev-kit node id). These two strings living in the same textual namespace
1164
- never implies a binding - a region id `"header"` and a target name
1165
- `"header"` bind to each other only because of an explicit declaration, not
1166
- because the strings match (verified by a dedicated test: the same
1167
- observation with and without the explicit declaration produces `bound`
1168
- only in the former case).
1169
-
1170
- ```ts
1171
- // domain/externalReferenceRuntimeBinding.ts
1172
- interface ReferenceRuntimeBindingDeclaration {
1173
- referenceRegion: string; // Prompt 2 ReferenceRegion.id
1174
- runtimeTarget: string; // v0.2 NamedTarget.name
1175
- }
1176
-
1177
- const REFERENCE_RUNTIME_BINDING_STATUSES = ['bound', 'ambiguous', 'unavailable'] as const;
1178
-
1179
- interface ReferenceRuntimeBindingResult {
1180
- referenceRegion: string;
1181
- runtimeTarget: string;
1182
- status: 'bound' | 'ambiguous' | 'unavailable';
1183
- reasonCode?: 'runtime-target-not-configured' | 'runtime-target-not-found' | 'runtime-target-ambiguous' | 'runtime-target-evidence-unavailable';
1184
- detail: string;
1185
- targetResolutionStatus?: TargetSelectionStatus; // v0.2's own resolution status, when evidence for it exists
1186
- targetVisible?: boolean; // provenance only - never affects status
1187
- }
1188
-
1189
- interface ReferenceRuntimeBindingEvaluation {
1190
- referenceId: string;
1191
- referenceRequestId: string;
1192
- candidateObservationId: string;
1193
- candidateRequestId: string;
1194
- compatibility: ComparabilityResult; // reused verbatim from v0.7 Prompt 4
1195
- bindings: ReferenceRuntimeBindingResult[]; // empty exactly when compatibility.state === 'incomparable'
1196
- }
1197
-
1198
- function evaluateReferenceRuntimeBindings(
1199
- reference: ExternalReferenceArtifact,
1200
- candidate: ObservationArtifact,
1201
- declarations: readonly ReferenceRuntimeBindingDeclaration[],
1202
- ): { ok: true; evaluation: ReferenceRuntimeBindingEvaluation } | { ok: false; reason: string };
1203
- ```
1204
-
1205
- Key rules:
1206
-
1207
- - **Explicit, never inferred.** A binding declaration is user/configuration
1208
- input asserting a conceptual correspondence; this module never discovers
1209
- it from screenshot geometry, matching names, matching text, or source
1210
- code. There is no automatic-matching algorithm anywhere in this module.
1211
- - **Reuses, never duplicates.** The compatibility gate reuses
1212
- `evaluateReferenceCandidateCompatibility` (Prompt 4) verbatim - viewport/
1213
- theme/application-state/authenticated-state comparison logic is never
1214
- re-implemented here. Runtime-target resolution reuses `targetPresence`
1215
- (v0.4 `comparisonEngine.ts`, now additively exported alongside
1216
- `assessOptionalComparabilityDimension`) - the exact same "how do I read a
1217
- `TargetEvidenceRecord`'s resolution" rule v0.4's own before/after target
1218
- comparison already uses. No second target resolver, no browser launch, no
1219
- Chromium query, no selector evaluation, no live-DOM inspection - this
1220
- module consumes only an already-captured `ObservationArtifact`'s
1221
- `requestConfig.targets`/`targetEvidence`.
1222
- - **Compatibility gates before evaluation, structurally.** If
1223
- `evaluateReferenceCandidateCompatibility` reports `incomparable`,
1224
- `bindings` is the empty array and the caller reads the reason from the
1225
- embedded `compatibility` field - there is no binding-local "incompatible"
1226
- status; Prompt 4's compatibility result is represented exactly once, not
1227
- duplicated into a parallel vocabulary.
1228
- - **Two-layer validation.** Reference-region existence, declaration shape,
1229
- bounds, and duplicate/conflict rules are validated structurally against
1230
- the `ExternalReferenceArtifact` alone (`isValidReferenceRuntimeBindingDeclarations`)
1231
- - independent of any candidate, mirroring Prompt 3's "unknown region
1232
- reference is a structural validation failure" precedent exactly: a
1233
- declaration naming a nonexistent reference region, or any reference with
1234
- no `regions` declared at all, fails the whole evaluation closed before a
1235
- candidate is even considered. Runtime-target availability, by contrast,
1236
- is evaluated per-candidate inside `evaluateReferenceRuntimeBindings`
1237
- itself, since the same declaration can be `bound` against one candidate
1238
- and `unavailable` against another.
1239
- - **Duplicate/conflicting-declaration rule** (mirrors Prompt 3's
1240
- requirement-subject uniqueness rule): no two declarations may name the
1241
- same `referenceRegion` (case-insensitively), whether they agree on
1242
- `runtimeTarget` (an exact duplicate) or disagree (a conflict) - both fail
1243
- the same way, never silently resolved by keeping the first. The reverse -
1244
- several distinct reference regions naming the same `runtimeTarget` - is
1245
- deliberately allowed (e.g. two design sub-regions legitimately
1246
- corresponding to one runtime container element).
1247
- - **Target-resolution-state handling.** `targetPresence`'s four outcomes
1248
- map onto binding status as: `matched` -> `bound`; `ambiguous` -> `ambiguous`
1249
- (the candidate's own configured target resolved ambiguously - never
1250
- reported bound even though its stable name exists); `not-found` ->
1251
- `unavailable` (`runtime-target-not-found` - the target was configured but
1252
- the resolver found nothing on the page); no usable resolution evidence at
1253
- all -> `unavailable` (`runtime-target-evidence-unavailable`). A declared
1254
- `runtimeTarget` that was never part of the candidate's configured target
1255
- set at all is a fifth, CLI/config-boundary-only outcome -> `unavailable`
1256
- (`runtime-target-not-configured`) - never a dynamic page search.
1257
- - **Hidden-target decision.** A uniquely resolved (`matched`) but hidden
1258
- target is still reported `bound` - visibility never changes `status`.
1259
- `targetVisible` (from the existing `TargetVisibility` evidence, when
1260
- available) is carried as provenance only. Binding identity (does a stable
1261
- correspondence exist) and later fidelity evaluability (can this evidence
1262
- actually be used to check the design) are treated as distinct questions;
1263
- this prompt answers only the former.
1264
- - **Not every region needs a binding.** `isValidReferenceRuntimeBindingDeclarations`
1265
- never requires full region coverage - a reference may have regions no
1266
- declaration names at all (they simply have no bound runtime target for
1267
- this candidate). This is not a completeness gate; Prompt 6 (or later) may
1268
- add one for the regions that selected requirements actually need.
1269
- - **No persisted artifact family.** `evaluateReferenceRuntimeBindings` is a
1270
- pure, on-demand function over an already-persisted reference, an
1271
- already-persisted candidate observation, and an in-memory declaration
1272
- collection. No `ExternalReferenceBindingArtifact` (or equivalent) is
1273
- introduced - the same "cheap to recompute, persisting invites drift"
1274
- reasoning Prompt 4 already applied to its own compatibility result.
1275
- Neither the reference nor the observation artifact is ever rewritten to
1276
- carry a binding result: a design reference may later be evaluated against
1277
- several different candidates, and one observation may be evaluated
1278
- against several different references, so binding is kept as downstream,
1279
- candidate-specific, reference-specific derived evidence rather than
1280
- mutating either immutable source artifact.
1281
- - **No new identity function.** Unlike `buildRequestIdentity`/
1282
- `buildExternalReferenceRequestIdentity`, no hash-based logical identity is
1283
- computed for a binding declaration or its evaluated result - there is no
1284
- persistence and no cross-document reference-by-id need yet (mirroring
1285
- Prompt 4's `ReferenceCandidateCompatibilityResult`, which took the same
1286
- approach). Provenance is instead carried directly as plain fields
1287
- (`referenceId`, `referenceRequestId`, `candidateObservationId`,
1288
- `candidateRequestId`, plus each result's own `referenceRegion`/
1289
- `runtimeTarget`) - already deterministic, already sufficient for a caller
1290
- to trace every result back to its inputs, without inventing a fifth
1291
- identity-hashing convention for a value this prompt does not persist.
1292
- - **Deterministic ordering.** `bindings` preserves authored declaration
1293
- order (mirroring the "authored order is semantic" convention already used
1294
- for regions/requirements) rather than sorting by any derived key.
1295
- - **Bounded.** `MAX_REFERENCE_RUNTIME_BINDINGS` (20) caps the declaration
1296
- collection, mirroring `MAX_REFERENCE_REGIONS`.
1297
- - **No public CLI surface yet.** Only the programmatic
1298
- `evaluateReferenceRuntimeBindings`/`isValidReferenceRuntimeBindingDeclarations`
1299
- functions are exported. A standalone CLI command was deliberately not
1300
- added merely for symmetry with `import-reference`/`observe`; Prompt 6
1301
- (structured fidelity evaluation) is expected to become the first concrete
1302
- consumer and public-surface owner for this capability.
1303
-
1304
- ## v0.7 Prompt 6 structured reference-vs-candidate fidelity evaluation
1305
-
1306
- Released as `0.7.0`. One new pure domain module,
1307
- `domain/externalReferenceFidelity.ts`, and its CLI-facing counterpart,
1308
- `application/referenceFidelityEvaluationService.ts` plus the new
1309
- `evaluate-reference-fidelity` CLI command - the first point in this whole
1310
- v0.7 stack where a reference's authored expectation is actually compared
1311
- against live candidate evidence. No new artifact field, no schema version
1312
- bump: this prompt reuses Prompt 1-5's artifacts and result types entirely.
1313
-
1314
- ```ts
1315
- // domain/externalReferenceFidelity.ts
1316
- const REFERENCE_REQUIREMENT_FIDELITY_STATUSES = ['pass', 'fail', 'unavailable'] as const;
1317
- const REFERENCE_FIDELITY_STATES = ['not-evaluated', 'pass', 'fail'] as const;
1318
- const REFERENCE_FIDELITY_BLOCK_REASONS = ['reference-inadequate', 'incompatible'] as const;
1319
-
1320
- interface ReferenceRequirementFidelityResult {
1321
- requirementId: string;
1322
- category: AuthoredChangeScopeCategory;
1323
- expectedDependentMode?: ExpectedDependentMode;
1324
- subject: ReferenceRequirementSubject;
1325
- boundRuntimeTargets: string[];
1326
- status: 'pass' | 'fail' | 'unavailable';
1327
- reasonCode?: 'reference-evidence-unavailable' | 'reference-relationship-not-exhibited' | 'binding-unavailable' | 'candidate-evidence-unavailable' | 'coordinate-mapping-unavailable'; // present iff status === 'unavailable'
1328
- detail?: string;
1329
- // region-property/region-measurement subjects only:
1330
- referenceValue?: number; // reference-image pixels
1331
- candidateRawValue?: number; // CSS pixels, as captured
1332
- candidateValue?: number; // candidateRawValue converted into reference-image-pixel space
1333
- delta?: number; // candidateValue - referenceValue
1334
- tolerance?: ReferenceRequirementTolerance;
1335
- // region-relationship subjects only:
1336
- expectedRelationship?: PairwiseRelationshipKind;
1337
- actualRelationship?: PairwiseRelationshipKind;
1338
- }
1339
-
1340
- interface ReferenceCandidateFidelityEvaluation {
1341
- referenceId: string;
1342
- referenceRequestId: string;
1343
- candidateObservationId: string;
1344
- candidateRequestId: string;
1345
- adequacy: ReferenceRequirementAdequacy; // reused verbatim from Prompt 3
1346
- compatibility?: ComparabilityResult; // reused verbatim from Prompt 4; absent only when adequacy itself is inadequate
1347
- bindings?: ReferenceRuntimeBindingEvaluation; // reused verbatim from Prompt 5; absent when an earlier gate blocked
1348
- state: 'not-evaluated' | 'pass' | 'fail';
1349
- blockedBy?: 'reference-inadequate' | 'incompatible'; // present iff state === 'not-evaluated'
1350
- requirementResults: ReferenceRequirementFidelityResult[]; // empty iff state === 'not-evaluated'
1351
- }
1352
-
1353
- function evaluateReferenceCandidateFidelity(
1354
- reference: ExternalReferenceArtifact,
1355
- candidate: ObservationArtifact,
1356
- bindingDeclarations: readonly ReferenceRuntimeBindingDeclaration[],
1357
- options?: { geometryTolerancePx?: number },
1358
- ): { ok: true; evaluation: ReferenceCandidateFidelityEvaluation } | { ok: false; reason: string };
1359
- ```
1360
-
1361
- **Result vocabulary.** `pass`/`fail`/`unavailable` is reused from v0.5's
1362
- `CLAUSE_RESULT_STATUSES` shape (the same honest three-state idea: a result
1363
- either satisfies its condition, fails it, or cannot be evaluated - never a
1364
- score) but is its own independently-owned constant, deliberately excluding
1365
- v0.5's fourth member, `'conflict'` - Prompt 6 has no cross-requirement
1366
- authoring-conflict concept (each requirement is evaluated independently
1367
- against its own subject), so reusing `conflict` would invite a status this
1368
- prompt can never actually produce.
1369
-
1370
- **Evaluation order (frozen, never reordered):** reference structural
1371
- validation -> candidate structural validation -> binding-declaration
1372
- structural validation -> Prompt 3 reference adequacy -> Prompt 4
1373
- compatibility -> Prompt 5 binding evaluation -> per-requirement candidate-
1374
- evidence/coordinate-mapping checks -> per-requirement tolerance/
1375
- relationship comparison -> overall result. The first three (structural
1376
- validation) failures return `{ ok: false, reason }` - a caller/config error,
1377
- never a fidelity outcome. The next two (adequacy `inadequate`, compatibility
1378
- `incomparable`) short-circuit to `state: 'not-evaluated'` with an empty
1379
- `requirementResults` - an earlier blocking gate never lets an ordinary
1380
- PASS/FAIL requirement set get fabricated past it. `adequacy` "partial" (some,
1381
- but not all, authored requirements individually unavailable) does **not**
1382
- block evaluation - it proceeds normally, and the individual unavailable
1383
- reference-side requirements simply also report `unavailable` at the
1384
- per-requirement level (their own reference-evidence problem, re-derived
1385
- identically by `evaluateOneRequirement`, not looked up from the adequacy
1386
- result).
1387
-
1388
- **Coordinate mapping - the central problem this prompt solves.** Prompt 2
1389
- regions and Prompt 3 tolerances are authored in reference-image pixels;
1390
- `ObservationArtifact` target geometry is CSS pixels. This module establishes
1391
- exactly one explicit, deterministic scale from `reference.applicability.viewport`
1392
- (the CSS-pixel runtime viewport, Prompt 4) and the reference image's own
1393
- pixel dimensions (Prompt 1) - `scaleX = imageWidth / viewportWidth`,
1394
- `scaleY = imageHeight / viewportHeight` - and converts every candidate
1395
- measurement into reference-image-pixel space before comparing it against a
1396
- Prompt 3 tolerance. It never assumes 1 reference-image pixel equals 1 CSS
1397
- pixel, and it never performs cropping, offset, rotation, or perspective
1398
- registration - only a deliberately bounded full-frame mapping. `scaleX`/
1399
- `scaleY` must agree within a small, independently-owned coordinate-mapping-
1400
- validity tolerance (1% relative, never a user-authored design tolerance) or
1401
- the mapping is rejected outright; a reference with no applicable viewport at
1402
- all likewise has no mapping. Either way, every numeric (`region-property`/
1403
- `region-measurement`) requirement becomes `unavailable`/`coordinate-mapping-unavailable`
1404
- - categorical `region-relationship` requirements are unaffected (they never
1405
- need a scale). Horizontal fields (`x`/`width`/`right`/`centerX` and the
1406
- horizontal measurements) always scale by `scaleX`; vertical fields (`y`/
1407
- `height`/`bottom`/`centerY` and the vertical measurements) always scale by
1408
- `scaleY` - this falls out automatically from converting a full
1409
- `TargetGeometry` into a `ReferenceRegionGeometry`-shaped value per axis,
1410
- never a hand-picked per-property axis table.
1411
-
1412
- **Tolerance is reused exactly, never redefined.** A single rule -
1413
- `abs(delta) <= allowedAmount` - covers all three Prompt 3 tolerance kinds:
1414
- `exact` is simply the zero-tolerance case (`allowedAmount = 0`);
1415
- `absolute-reference-px` uses its authored `amount` directly (already in
1416
- reference-image pixels); `percent`'s denominator is `Math.abs(referenceValue)`,
1417
- mirroring v0.5's own `toleranceToPx` "may vary by up to N%" convention
1418
- exactly (independently reimplemented in reference-image-pixel units, never
1419
- imported - `frontendContractEvaluation.ts`'s `ContractTolerance` is a
1420
- different, CSS-pixel-implicit unit). No hidden epsilon is added anywhere;
1421
- subpixel precision is preserved through to the final comparison, so a
1422
- tolerance-boundary value (e.g. delta exactly equal to the allowed amount)
1423
- passes and one unit past it fails, exactly as authored.
1424
-
1425
- **Region-property evaluation** reads `TargetGeometry` from the bound
1426
- target's `targetEvidence` entry, converts it into a `ReferenceRegionGeometry`-
1427
- shaped value (adding `centerX`/`centerY`, computed identically to
1428
- `deriveReferenceRegionGeometry`) both raw (CSS) and scaled (reference-image
1429
- pixels), and reads `[subject.property]` off each - `candidateRawValue`
1430
- (CSS) and `candidateValue` (reference-image pixels) are both reported.
1431
-
1432
- **Region-measurement evaluation** converts *both* bound targets' geometries
1433
- the same way and calls the existing `deriveReferenceRequirementMeasurement`
1434
- (Prompt 3) on the converted geometries directly - reusing Prompt 3's exact
1435
- gap/delta formulas rather than reimplementing a parallel "runtime version"
1436
- of them, and never inventing a generic geometry expression language. A
1437
- geometrically-undefined gap (the two targets overlap on the relevant axis)
1438
- is `unavailable`, mirroring Prompt 3's own reference-side treatment of the
1439
- identical situation.
1440
-
1441
- **Region-relationship evaluation** first confirms the reference itself
1442
- actually exhibits its own selected relationship (`deriveReferenceRequirementExpectation`'s
1443
- `matches` field) - if not, the result is `unavailable`/
1444
- `reference-relationship-not-exhibited` (a reference-authoring problem, never
1445
- a candidate `fail`). It then resolves both bound targets and calls the
1446
- canonical `deriveLayoutRelationships` (v0.4) over the *whole* candidate
1447
- observation - never a second, parallel relationship formula - and looks up
1448
- the pairwise record for the bound target pair **scoped to the exact
1449
- requested relationship family** (an independently-owned, third duplicate of
1450
- the same `RELATIONSHIP_FAMILY_GROUPS` shape already used by
1451
- `frontendContractEvaluation.ts` and `externalReferenceRequirements.ts` -
1452
- this is the same real bug class Prompt 3 fixed: matching the first record
1453
- for a target pair regardless of family would silently compare against the
1454
- wrong relationship kind). A record only derivable in the reversed target
1455
- order is `unavailable`, never auto-flipped - identical to Prompt 3's own
1456
- reference-side handling of the same situation. `pass` requires the
1457
- candidate's actual relationship kind to equal the requirement's authored
1458
- `relationship` exactly.
1459
-
1460
- **Binding gate.** Every subject's dependent reference region(s) must have a
1461
- `bound` (never `ambiguous`/`unavailable`, and never simply absent from the
1462
- supplied declarations) Prompt 5 binding result, or the requirement is
1463
- `unavailable`/`binding-unavailable` - this module never guesses another
1464
- target and never auto-binds based on geometry or names. A `bound` target
1465
- that is not visible (`TargetVisibility.visible !== true`, including when
1466
- visibility evidence itself is unavailable) is treated as having no usable
1467
- geometry - `unavailable`/`candidate-evidence-unavailable` - preserving the
1468
- distinction between "binding succeeded" (Prompt 5's question) and "this
1469
- evidence is usable for fidelity evaluation" (this prompt's question): a
1470
- hidden-but-uniquely-resolved target still has a stable identity, but its
1471
- geometry is never treated as meaningful for a numeric/relationship
1472
- comparison.
1473
-
1474
- **Categories are preserved, never given different PASS/FAIL rules.**
1475
- `category`/`expectedDependentMode` are carried through to each result as
1476
- provenance only; Prompt 3 never implemented a `required`-vs-`permitted`
1477
- directional evaluation difference for its own expectation/adequacy
1478
- derivation (unlike v0.5's runtime-directional contract clauses), so Prompt 6
1479
- does not invent one now - every requirement in the reference's authored
1480
- collection is evaluated by the identical rule and counts identically toward
1481
- the overall result, regardless of category.
1482
-
1483
- **Overall fidelity result.** For an evaluated (non-blocked) pair, `state`
1484
- is `'pass'` only when every requirement result is `'pass'`; any `'fail'` or
1485
- `'unavailable'` result forces `state: 'fail'` - there is no meaningful third
1486
- overall bucket once evaluation has actually run, since "some/all
1487
- unavailable" and "some/all fail" both equally mean "not every selected
1488
- requirement is confirmed satisfied." A reference with zero selected
1489
- requirements never reaches this stage at all - it is `inadequate` (Prompt
1490
- 3's own zero-requirements rule) and therefore `not-evaluated`, never a
1491
- meaningless `pass`.
1492
-
1493
- **Persistence decision: none.** `evaluateReferenceCandidateFidelity` (and
1494
- its CLI-facing wrapper, `evaluateReferenceCandidateFidelityFromArtifactRoots`)
1495
- is a pure, on-demand function over already-persisted/in-memory evidence -
1496
- no new `ExternalReferenceFidelityEvaluationArtifact` (or equivalent) is
1497
- introduced. Rationale, identical to Prompt 4/5's own precedent: the result
1498
- is cheap to recompute deterministically from its inputs (a reference, a
1499
- candidate, and a caller-supplied binding-declaration collection), and
1500
- persisting it would invite drift with no corresponding benefit at this
1501
- stage; this may be revisited only if Prompt 7's architecture proves
1502
- persistence necessary.
1503
-
1504
- **CLI**: `evaluate-reference-fidelity --reference <root> --candidate <root>
1505
- [--bindings-file <json-file>] [--enforce]` - the CLI surface Prompt 5
1506
- deliberately deferred. `--bindings-file` follows the exact
1507
- `--requirements-file`/`--regions-file` wrapped-object convention
1508
- (`{ "bindings": [...] }`); CLI code owns only flag syntax/file reading/JSON
1509
- parsing/root-shape validation, with every binding-declaration rule staying
1510
- owned by `isValidReferenceRuntimeBindingDeclarations`. `--enforce` mirrors
1511
- `evaluate-contract`'s exact precedent: it changes only the process exit
1512
- status for an already-computed `state: 'fail'` result, never the printed
1513
- content - and has no effect on `not-evaluated`, which always exits 0 (a
1514
- compatibility/adequacy blocker is a successful, structured, honest
1515
- non-evaluation, never an execution error and never a design mismatch).
1516
- Persists nothing; there is no `--output` flag.
1517
-
1518
- ## v0.7 Prompt 7 bounded reference-fidelity projection and v0.6 bounded-agent-context integration
1519
-
1520
- Released as `0.7.0`. Additive extension of the v0.6 bounded-agent-context
1521
- contract above and of the v0.7 Prompt 6 fidelity evaluator - no new bounded-
1522
- context artifact family, no second visual-context system, no schema version
1523
- bump (`BOUNDED_AGENT_CONTEXT_SCHEMA_VERSION` stays `1.0.0`, following the
1524
- exact precedent already set when `correlations?` was added in v0.6 Batch 3).
1525
-
1526
- **Chosen integration owner.** `projectBoundedAgentContext` itself gains one
1527
- new optional input (`fidelity?: ReferenceCandidateFidelityEvaluation`, plus
1528
- `fidelityRequired?: boolean`) rather than a separate `VisualAgentContext`/
1529
- `VisualPromptPacket`/`ReferencePromptBuilder`. This was chosen over a pure
1530
- post-hoc "attach" step (the shape `attachRuntimeStaticCorrelations` uses)
1531
- because fidelity-relevant runtime targets must compete fairly for
1532
- `MAX_RUNTIME_TARGETS` capacity and receive the exact same geometry/
1533
- visibility/screenshot assembly contract-clause-derived targets already get -
1534
- an attach-only step run after target allocation could never produce that. A
1535
- new pure module, `domain/referenceFidelityProjection.ts`
1536
- (`projectReferenceFidelity`), derives the bounded, prioritized fidelity
1537
- content plus the target-id/omission/truncation contributions
1538
- `projectBoundedAgentContext` folds into its own existing pipeline - it is
1539
- not a second fidelity-evaluation engine, only a selection over Prompt 6's
1540
- already-computed result.
1541
-
1542
- ```ts
1543
- // domain/boundedAgentContext.ts - additive
1544
- interface BoundedAgentContextSourceReferences {
1545
- // ...unchanged fields...
1546
- referenceId?: string; // new, optional
1547
- referenceRequestId?: string; // new, optional
1548
- }
1549
-
1550
- const MAX_FIDELITY_MISMATCHES = 15;
1551
- const MAX_FIDELITY_PROTECTED_CONTEXT = 10; // reuses MAX_RELATIONSHIP_EVIDENCE_PER_TARGET's value
1552
-
1553
- interface BoundedReferenceFidelityProjection {
1554
- referenceId: string;
1555
- referenceRequestId: string;
1556
- candidateObservationId: string;
1557
- candidateRequestId: string;
1558
- adequacy: ReferenceRequirementAdequacy; // reused verbatim from Prompt 3
1559
- compatibility?: ComparabilityResult; // reused verbatim from Prompt 4
1560
- state: ReferenceFidelityState; // reused verbatim from Prompt 6
1561
- blockedBy?: ReferenceFidelityBlockReason; // reused verbatim from Prompt 6
1562
- mismatches: ReferenceRequirementFidelityResult[]; // bounded, prioritized non-pass requirements (Prompt 6 type, unmodified)
1563
- protectedContext: ReferenceRequirementFidelityResult[]; // bounded passing protected/preserved requirements, as "do not break this" context
1564
- }
1565
-
1566
- interface BoundedAgentContextArtifact {
1567
- // ...unchanged fields...
1568
- fidelity?: BoundedReferenceFidelityProjection; // new, optional - mirrors `correlations?`'s own additive precedent exactly
1569
- }
1570
- ```
1571
-
1572
- **Selection policy** (`domain/referenceFidelityProjection.ts#projectReferenceFidelity`):
1573
- only Prompt 6's non-`pass` requirement results are ever candidates for
1574
- `mismatches` - passing requirements are never dumped by default, satisfying
1575
- this prompt's "bounded coding-agent use" design goal. Each candidate is
1576
- classified into a tier by its authored category/mode, reusing
1577
- `boundedAgentContextProjection.ts#clauseTier`'s exact rule (duplicated, not
1578
- imported, per this repository's established per-module small-helper
1579
- convention - never a reference-specific protected/preserved taxonomy):
1580
- `protected`/`preserved` are always `required`; `expected-dependent` is
1581
- `required` only in `'required'` mode; `requested` and `expected-dependent`/
1582
- `'permitted'` are `optional`.
1583
-
1584
- **Priority policy**: 1) `fail` + `required` tier, 2) `unavailable` +
1585
- `required` tier, 3) any other non-`pass` (optional-tier) result. Within one
1586
- priority class, Prompt 6's own authored requirement order is preserved (a
1587
- stable sort by priority rank only) - never re-ranked by an opaque score.
1588
- The final `mismatches` array is reported in priority order (highest first),
1589
- not restored to authored order, since the whole point of prioritization is
1590
- that the most actionable evidence appears first when the set is large.
1591
-
1592
- **Cap values**: `MAX_FIDELITY_MISMATCHES = 15` and
1593
- `MAX_FIDELITY_PROTECTED_CONTEXT = 10` (reusing
1594
- `MAX_RELATIONSHIP_EVIDENCE_PER_TARGET`'s value) - both judgment-call bounds
1595
- in the same spirit as v0.6 Batch 1's own frozen caps (no measured fixture
1596
- corpus exists yet for either concept).
1597
-
1598
- **Omission/truncation behavior**: reuses `OmissionRecord`/`TruncationRecord`
1599
- wholesale, no second reporting model. When mismatches exceed the cap, a
1600
- `{subject: 'fidelity-mismatches', limit, actualCount, required}` truncation
1601
- is recorded, plus one `{subject: 'fidelity-mismatch:<requirementId>',
1602
- reason: 'required-evidence-lost-by-bound', required: true}` omission for
1603
- *each* dropped required-tier mismatch (optional-tier drops are truncated
1604
- but never separately omitted as "required loss", since they were never
1605
- required). `protectedContext` truncation is always `required: false` - it
1606
- is confirmatory/passing context, never a design-fidelity failure. These
1607
- records are folded into `projectBoundedAgentContext`'s own `omissions`/
1608
- `truncations` arrays *before* its existing aggregate `capOmissions`/
1609
- `capTruncations` calls and its existing adequacy computation run - fidelity
1610
- loss is never a separate adequacy code path, it simply participates in the
1611
- exact same `anyRequiredLoss`/`anyOptionalLoss` rule every other evidence
1612
- source already uses.
1613
-
1614
- **Adequacy behavior**: a `not-evaluated` fidelity (blocked by Prompt 6's own
1615
- `reference-inadequate`/`incompatible` gates) is never converted into "no
1616
- problems" - `projectReferenceFidelity` records an explicit
1617
- `{subject: 'fidelity', reason: 'unsupported-or-unavailable', required,
1618
- detail}` omission, where `required` defaults to `true` (supplying a
1619
- fidelity evaluation to be projected at all is itself the signal that the
1620
- task depends on it, mirroring `CorrelationTargetInput.required`'s existing
1621
- v0.6 convention - callers who want fidelity as purely incidental context set
1622
- `fidelityRequired: false`). A `required: true` fidelity omission, folded
1623
- into the existing adequacy computation, prevents `adequacy.state` from
1624
- remaining `'adequate'` (it becomes `'partial'`, or `'inadequate'` when
1625
- combined with other required loss reaching the existing threshold) - it is
1626
- never silently ignored. A `required: false` omission can degrade adequacy
1627
- to at most `'partial'`, per v0.6's own pre-existing "optional-only loss
1628
- never means inadequate" rule - unchanged, not redefined. A `pass` fidelity
1629
- result contributes no omissions/truncations at all and never degrades
1630
- adequacy.
1631
-
1632
- **Not-evaluated fidelity behavior**: preserved exactly as Prompt 6 reported
1633
- it - `fidelity.state`/`fidelity.blockedBy` on the output artifact are a
1634
- direct pass-through of Prompt 6's own values, with `mismatches`/
1635
- `protectedContext` both empty (there is nothing to select from an empty
1636
- `requirementResults`).
1637
-
1638
- **Per-target organization**: every fidelity mismatch's `boundRuntimeTargets`
1639
- (Prompt 5/6's own field, never truncated) becomes a required- or permitted-
1640
- tier addition to `projectBoundedAgentContext`'s existing target-id sets,
1641
- so those runtime targets receive full `BoundedRuntimeTargetProjection`
1642
- treatment (geometry/visibility/overflow/scrollOwner/screenshotRef) through
1643
- the exact existing assembly code - no duplicated target-projection logic.
1644
- Reference regions are never used as a correlation or target-selection key;
1645
- only the already-bound stable v0.2 runtime target ids are.
1646
-
1647
- **Multi-target relationship representation**: a `region-relationship`
1648
- mismatch's `boundRuntimeTargets` array (already carrying both bound
1649
- targets, from Prompt 6) is used as-is - both targets are added to the
1650
- required/permitted set, so both appear in `targets`. Nothing collapses a
1651
- two-target relationship failure onto a single target.
1652
-
1653
- **Static-correlation reuse**: entirely unchanged. `deriveRuntimeStaticCorrelations`/
1654
- `attachRuntimeStaticCorrelations` are not modified, not called from within
1655
- this prompt's new code, and remain the caller's own separate step -
1656
- `BoundedRuntimeTargetProjection.targetId`/`RuntimeStaticCorrelationRecord.runtimeTargetId`
1657
- already share the same stable v0.2 identity a fidelity mismatch's
1658
- `boundRuntimeTargets` also uses, so a caller (Prompt 8) joins fidelity,
1659
- target, and correlation evidence by that one shared id without this module
1660
- ever needing to read source, run my-dev-kit, or choose among ambiguous
1661
- candidates itself.
1662
-
1663
- **Ambiguous/unavailable correlation behavior**: unaffected - a
1664
- `RuntimeStaticCorrelationRecord` with `status: 'ambiguous'` continues to
1665
- preserve every competing candidate (v0.6's own frozen invariant,
1666
- untouched), and `status: 'unavailable'` never causes a fidelity mismatch
1667
- for that same runtime target to be dropped - the two evidence kinds
1668
- (runtime fidelity, static correlation) are attached independently and
1669
- neither erases the other.
1670
-
1671
- **Provenance**: every included mismatch remains traceable to the reference
1672
- (`sources.referenceId`/`referenceRequestId`, new), the requirement
1673
- (`requirementId`, `category`, `subject` - naming its reference region(s)),
1674
- the Prompt 5 binding (`boundRuntimeTargets`), the candidate
1675
- (`sources.observationIds`), and the full Prompt 6 evidence
1676
- (`referenceValue`/`candidateRawValue`/`candidateValue`/`delta`/`tolerance`
1677
- or `expectedRelationship`/`actualRelationship`) - nothing is replaced by a
1678
- prose-only summary. No raw image bytes are ever embedded (fidelity carries
1679
- only identifiers and numeric/categorical evidence, never pixels), and no
1680
- source-ownership field (`sourceOwner`/`sourceFile`/`component`/`symbol`/
1681
- `causedBy`) is ever produced - Prompt 7 stops at the runtime target exactly
1682
- as Prompt 6 did; v0.6's own, unmodified static correlation is the only
1683
- source-adjacent evidence this context ever carries, and it remains
1684
- evidence, never edit authorization.
1685
-
1686
- **Identity impact**: `buildBoundedAgentContextRequestIdentity` gained a
1687
- final optional `fidelity?: unknown` parameter - omitted (never `null`) from
1688
- the hashed semantic view when absent, so every pre-Prompt-7 call site keeps
1689
- producing its exact byte-identical hash (verified by a frozen-vector-style
1690
- regression test). When present, the caller's already-derived, bounded
1691
- `BoundedReferenceFidelityProjection` (not the raw Prompt 6 evaluation) is
1692
- hashed, so identity changes exactly when the content a caller would
1693
- actually receive changes - never merely because an unselected, dropped
1694
- requirement result changed somewhere upstream. `sources.referenceId`/
1695
- `referenceRequestId` follow the identical omit-when-absent convention.
1696
- Operational paths were never an identity input for this artifact family to
1697
- begin with (no path parameter exists anywhere in this contract), so path
1698
- independence holds trivially.
1699
-
1700
- **Schema-version decision**: no bump. Every new field
1701
- (`BoundedAgentContextSourceReferences.referenceId`/`referenceRequestId`,
1702
- `BoundedAgentContextArtifact.fidelity`) is additive and optional; a
1703
- pre-Prompt-7 artifact/consumer remains fully valid and behaviorally
1704
- unchanged with all of them absent, matching the exact precedent
1705
- `correlations?` already established without a version bump in v0.6 Batch 3.
1706
-
1707
- **Persistence decision: none.** `projectBoundedAgentContext` and
1708
- `projectReferenceFidelity` both remain pure, programmatic, in-memory
1709
- functions - no new writer/reader, no new artifact family. This mirrors
1710
- Prompt 6's own "no persisted fidelity artifact" decision and v0.6's
1711
- existing "bounded agent context is library-only" architecture.
1712
-
1713
- **CLI decision**: none added. v0.6 bounded agent context has never had a
1714
- CLI surface, and this prompt does not introduce one - Prompt 8 is expected
1715
- to become the first concrete consumer of `projectBoundedAgentContext`'s
1716
- (now fidelity-aware) programmatic output.
1717
-
1718
- ## v0.7 Prompt 8 controlled end-to-end external-reference coding-agent correction workflow
1719
-
1720
- Released as `0.7.0`. One new pure domain module,
1721
- `domain/referenceCorrectionWorkflow.ts`, plus its identity counterpart,
1722
- `domain/referenceCorrectionIdentity.ts` - the first stage that composes
1723
- every Prompt 1-7 and v0.1/v0.4/v0.5/v0.6 owner into one traceable
1724
- reference-driven correction cycle. It reimplements none of them: reference
1725
- lifecycle/adequacy (Prompt 1/3), compatibility (Prompt 4), binding (Prompt
1726
- 5), fidelity (Prompt 6), bounded context (Prompt 7), runtime comparison
1727
- (v0.4 `compareObservations`), and contract evaluation (v0.5
1728
- `evaluateFrontendContract`) are all called, never re-derived. No new
1729
- persisted artifact family, no CLI surface, no remote AI dependency, and no
1730
- mechanism anywhere in this module (or any module it calls) that edits
1731
- target source.
1732
-
1733
- ```ts
1734
- // domain/referenceCorrectionWorkflow.ts
1735
- function prepareReferenceCorrection(input: {
1736
- reference: ExternalReferenceArtifact; // must be approved
1737
- baselineObservation: ObservationArtifact; // approved baseline / pre-change state
1738
- baselineContract: PersistentBaselineContract;
1739
- changeContract: PerChangeContract;
1740
- bindingDeclarations: readonly ReferenceRuntimeBindingDeclaration[];
1741
- currentObservation: ObservationArtifact; // fidelity is measured against this (= baselineObservation for the canonical proof)
1742
- generatedAt: string; producerVersion: string; projectionProfile: ProjectionProfile;
1743
- }): { ok: true; status: 'handoff-ready'; reviewRequestId: string; fidelity: ReferenceCandidateFidelityEvaluation; handoff: ReferenceCorrectionHandoff }
1744
- | { ok: true; status: 'blocked-not-evaluated'; reviewRequestId: string; fidelity: ReferenceCandidateFidelityEvaluation }
1745
- | { ok: false; reason: string };
1746
-
1747
- function reviewReferenceCorrectionAttempt(input: {
1748
- // ...same reference/baselineObservation/baselineContract/changeContract/bindingDeclarations...
1749
- reviewRequestId: string; // must match the id prepareReferenceCorrection returned for this exact semantic review
1750
- candidateObservation: ObservationArtifact; // fresh, post-edit capture
1751
- priorAttemptId?: string;
1752
- }): { ok: true; attempt: ReferenceCorrectionAttemptResult } | { ok: false; reason: string };
1753
- ```
1754
-
1755
- **Workflow architecture.** A narrowly-scoped coordinator, not a second
1756
- workflow engine: it holds no stage catalog, no job scheduler, and no
1757
- generic orchestration graph. It performs exactly two operations - "prepare"
1758
- (pre-change evidence -> bounded handoff) and "review" (post-edit candidate
1759
- -> one composed overall result) - matching this prompt's own explicit
1760
- guidance that the external-edit boundary must remain a visible seam between
1761
- two separate calls, never one command that blocks waiting for an external
1762
- actor.
1763
-
1764
- **New owners introduced**: `prepareReferenceCorrection`,
1765
- `reviewReferenceCorrectionAttempt` (composition only - no new evaluation
1766
- logic), `buildReferenceCorrectionReviewIdentity`/
1767
- `buildReferenceCorrectionAttemptIdentity` (deterministic identity, see
1768
- below), and the plain `ReferenceCorrectionHandoff`/
1769
- `ReferenceCorrectionAttemptResult` result shapes.
1770
-
1771
- **Existing owners reused, verbatim**: `isApprovedExternalReferenceArtifact`
1772
- (Prompt 1), `isValidReferenceRuntimeBindingDeclarations` (Prompt 5),
1773
- `evaluateReferenceCandidateFidelity` (Prompt 6), `projectBoundedAgentContext`
1774
- (Prompt 7, itself now fidelity-aware), `compareObservations` (v0.4),
1775
- `evaluateFrontendContract` (v0.5). None of their internal logic is
1776
- inspected, duplicated, or reimplemented by this module - only their
1777
- top-level results are read.
1778
-
1779
- **Approved-reference/approved-baseline requirement.** `prepareReferenceCorrection`
1780
- and `reviewReferenceCorrectionAttempt` both fail closed (`{ok: false}`) if
1781
- `reference` is not in the `'approved'` lifecycle state (Prompt 1's own
1782
- `isApprovedExternalReferenceArtifact` guard) - an imported-but-unapproved
1783
- reference is never treated as an authoritative target design. Neither
1784
- function ever calls `approveExternalReference`/`approveAndPersistBaseline`
1785
- itself; approval remains the caller's own separate, explicit action.
1786
-
1787
- **Pre-change candidate = approved baseline observation**, for the canonical
1788
- proof: `prepareReferenceCorrection`'s `currentObservation` and
1789
- `baselineObservation` are the same value, so the initial reference fidelity
1790
- can genuinely `FAIL` (measuring the gap between the current, already-
1791
- approved implementation and the desired new design) while the baseline
1792
- itself stays fully valid and approved. A caller's own architecture may
1793
- supply a distinct `currentObservation` only when justified - the workflow
1794
- does not require them to be identical, only that `currentObservation` and
1795
- `candidateObservation` are always independently validated
1796
- `ObservationArtifact`s.
1797
-
1798
- **Preparation (phase A)**: validates the common preconditions (approved
1799
- reference, matching baseline/contract coherence, valid binding
1800
- declarations), evaluates reference fidelity via Prompt 6 against
1801
- `currentObservation`, and - only when that evaluation actually produced a
1802
- result (`state !== 'not-evaluated'`) - projects it into a bounded context
1803
- via Prompt 7/v0.6 and returns the `ReferenceCorrectionHandoff`. A
1804
- `not-evaluated` fidelity (inadequate reference, or reference/candidate
1805
- incompatible state) is reported as `status: 'blocked-not-evaluated'` -
1806
- carrying the full Prompt 6 result for inspection, but never a fabricated
1807
- handoff pretending evidence is adequate. An ambiguous or unavailable
1808
- required binding does **not** block preparation outright - it still
1809
- produces a `handoff-ready` result, with the ambiguity/unavailability
1810
- visible directly in that requirement's own `unavailable`/`binding-
1811
- unavailable` mismatch (Prompt 6's own honest per-requirement reporting,
1812
- unchanged), so the external actor sees exactly why that specific
1813
- requirement cannot yet be assessed.
1814
-
1815
- **The handoff** (`ReferenceCorrectionHandoff`) carries `reviewRequestId`,
1816
- `referenceId`/`referenceRequestId`, `baselineObservationId`,
1817
- `currentObservationId`, the full Prompt 7 `boundedContext` (already
1818
- containing bounded fidelity mismatches, protected/preserved context,
1819
- adequacy/omission/truncation, and - when the caller supplied it - runtime/
1820
- static correlation), and a fixed, four-line `verificationPlan` explaining
1821
- in plain language what will be re-checked after the edit (fresh Chromium
1822
- capture, reference re-evaluation, v0.4/v0.5 re-evaluation, and the exact
1823
- overall-PASS rule) - never reduced to "make it look like the screenshot".
1824
- No raw reference image bytes, no full `ObservationArtifact`, and no source
1825
- excerpt are ever included.
1826
-
1827
- **Handoff persistence: none.** The handoff is a plain, JSON-serializable,
1828
- in-memory value returned directly to the caller - no new writer/reader, no
1829
- new artifact family. A caller that needs the handoff to cross a process/
1830
- session boundary (e.g. to hand it to an external coding-agent process) is
1831
- free to serialize it with its own mechanism; observer product code does not
1832
- own a persisted handoff artifact. This was a deliberate "smallest possible"
1833
- choice: the handoff's only genuinely new identity is `reviewRequestId`
1834
- (already deterministic and recomputable from stable inputs - see below), so
1835
- nothing about it requires observer-managed persistence to remain
1836
- traceable.
1837
-
1838
- **Review identity** (`buildReferenceCorrectionReviewIdentity`): a pure
1839
- function of `{referenceRequestId, baselineObservationId,
1840
- baselineContractId, baselineContractClauses, changeContractId,
1841
- changeContractClauses, bindingDeclarations}` only - never a timestamp,
1842
- never an operational file path. Deliberately hashes each contract's own
1843
- authored `clauses` content, not merely its `baselineId`/`contractId` label:
1844
- unlike this repository's content-derived identities elsewhere (e.g.
1845
- `ObservationArtifact.observationId`), a `PersistentBaselineContract`'s
1846
- `baselineId` and a `PerChangeContract`'s `contractId` are plain, caller-
1847
- authored strings (`approveAndPersistBaseline` persists `contract.baselineId`
1848
- verbatim, never recomputing it from `clauses`) - so two structurally valid
1849
- contracts could in principle share an id while authoring different clauses.
1850
- Hashing clause content directly closes that gap (caught during this
1851
- prompt's own independent-judge review before being reported PASS - see the
1852
- report's Tooling incidents section). `reviewReferenceCorrectionAttempt`
1853
- recomputes this same hash from its own inputs and rejects the call
1854
- (`{ok: false}`) if the caller-supplied `reviewRequestId` does not match -
1855
- this is the mechanism that makes "no hidden baseline change" an enforced
1856
- invariant rather than a documented intention: an attempt claiming to belong
1857
- to a review while actually supplying a different baseline observation,
1858
- baseline contract (id or clause content), per-change contract (id or clause
1859
- content), reference, or binding set can never silently succeed.
1860
-
1861
- **Attempt identity** (`buildReferenceCorrectionAttemptIdentity`): a pure,
1862
- deterministic function of `{reviewRequestId, candidateObservationId}` only
1863
- - deliberately never a fresh random nonce. Every candidate observation
1864
- already carries its own fresh, collision-resistant instance identity (v0.1's
1865
- `buildObservationIdentity`), so hashing it together with the review it was
1866
- captured for gives an attempt id that is both reproducible (the same
1867
- review+candidate pair always yields the same `attemptId`) and guaranteed
1868
- distinct per real capture.
1869
-
1870
- **Attempt history**: caller-managed, not observer-persisted. Because both
1871
- workflow functions are pure (no internal mutable state, no side effects),
1872
- an already-returned `ReferenceCorrectionAttemptResult` can never be
1873
- overwritten by a later call - a caller that keeps every attempt result it
1874
- receives (in memory, in its own log, or in its own storage) has a complete,
1875
- immutable, traceable history for free, linked via each attempt's own
1876
- `reviewRequestId` (shared across all attempts of one review),
1877
- `priorAttemptId` (an optional, purely informational link to the immediately
1878
- preceding attempt, carried through unchanged - never consulted by the
1879
- evaluation logic itself), and `attemptId`.
1880
-
1881
- **Baseline-across-attempts rule**: enforced structurally, not merely
1882
- documented. Every call to `reviewReferenceCorrectionAttempt` requires the
1883
- caller to re-supply `baselineObservation`/`baselineContract` in full, and
1884
- `compareObservations`/`evaluateFrontendContract` are always invoked with
1885
- that same baseline against the fresh `candidateObservation` - there is no
1886
- code path anywhere in this module that compares one candidate against a
1887
- prior candidate instead. Combined with the `reviewRequestId` coherence
1888
- check above, a caller cannot silently swap in a different baseline between
1889
- attempts of the same review without the call being rejected.
1890
-
1891
- **Overall result composition.** `ReferenceCorrectionOverallState =
1892
- 'not-evaluated' | 'pass' | 'fail'`:
1893
-
1894
- - `fidelity.state === 'not-evaluated'` -> overall `'not-evaluated'` - Prompt
1895
- 6's own explicit blocked state is preserved exactly, never collapsed into
1896
- an ordinary `'fail'`.
1897
- - otherwise, `fidelity.state === 'pass' && contractEvaluation.overallVerdict === 'PASS'`
1898
- -> overall `'pass'`; anything else -> overall `'fail'`.
1899
-
1900
- A structurally-incomparable baseline/candidate pair is *not* given its own
1901
- third overall bucket - v0.5's own `evaluateFrontendContract` already
1902
- returns `'FAIL'` (never `'PASS'`) for that case, per its own established,
1903
- unmodified precedent, and this workflow reuses that decision rather than
1904
- re-litigating it. `approvalEligible` is a plain, read-only boolean
1905
- (`true` iff `overallState === 'pass'`) - it is never itself an approval
1906
- action; the caller must still invoke the existing explicit
1907
- `approveAndPersistBaseline`/`approveExternalReference` owners separately,
1908
- and neither is ever called from within this module.
1909
-
1910
- **Correction iteration**: `reviewReferenceCorrectionAttempt` is called once
1911
- per candidate; the caller decides whether and when to call it again after
1912
- another external edit. There is no loop, no polling, no automatic retry,
1913
- and no mechanism in this module that itself waits for or drives an external
1914
- implementation step - the production boundary between "prepare a handoff"
1915
- and "review a candidate" is the explicit seam a human or an external
1916
- process controls.
1917
-
1918
- **Source-editing boundary**: absolute. Neither this module nor anything it
1919
- calls opens, reads, parses, or writes any target source file; both public
1920
- operations accept only already-captured `ObservationArtifact`s and already-
1921
- approved contract/reference artifacts. Real-Chromium candidate capture is
1922
- always the caller's own responsibility, through the existing, unmodified
1923
- observation pipeline (`runBrowserCapture`/`buildObservationArtifact`, the
1924
- same functions `application/observationPersistence.ts#observe` already
1925
- uses) - Prompt 8 adds no second browser adapter, screenshot engine, target
1926
- resolver, or evidence-capture path.
1
+ # Contracts
2
+
3
+ ## Visual-change workflow artifact (`1.0.0`, v0.10 Batch 1)
4
+
5
+ `my-frontend-observer/visual-change-workflow` is the immutable Observer-owned history envelope for one frozen visual-change request. It records a deterministic `visualChangeRequestId`, a fresh `visualChangeWorkflowId`, optional forward-only `supersedesVisualChangeWorkflowId`, exact references to existing canonical evidence, zero to twenty bounded attempt records, optional activation/governance result references, producer metadata, and creation provenance.
6
+
7
+ The two entry modes are exactly `actual-frontend` and `reference`. Reference mode stores validated explicit runtime binding declarations in canonical `bindings.json`; the manifest pins its SHA-256 and declaration count. Actual mode never writes that file. Project-relative evidence references must remain contained portable paths; source artifacts and media are referenced rather than copied.
8
+
9
+ Request identity hashes semantic scope only. It excludes timestamps and operational storage locations. Workflow instance identity is fresh for every explicit persistence. Attempt identity is deterministic over the visual-change request ID and candidate observation ID. Human review state is exactly `pending`, `correction-requested`, `accepted`, or `abandoned`; this foundation validates structure but does not implement the later acceptance rule or execute checks.
10
+
11
+ The v0.10 Batch 2 application composition explicitly activates either the frozen per-change contract or the frozen approved reference and workflow-owned bindings in project acceptance. Activation and restoration preserve unrelated configuration, use compensating atomic writes across mutable project configuration and immutable workflow revisions, and refuse acceptance drift. A workflow check resolves an alias only when its catalog identity and artifact location exactly match the frozen baseline, then invokes the existing `checkProject` owner once. Only a canonical result containing both baseline and candidate summaries can append a pending attempt revision; the snapshot is a bounded direct projection of that result and does not repeat evaluation.
12
+
13
+ The v0.10 Batch 3 Viewer exposes bounded workflow inspection at `GET /api/visual-changes/:handle/view` in both standalone and project-aware sessions. Create, activate, check, and restore POST operations are project-aware only, reuse the existing in-memory authoring capability and request guards, accept no filesystem paths, and delegate to the Batch 2 application service. Every new response is `no-store`. Successful immutable mutations return the exact new workflow ID so the Viewer can refresh and reselect that revision without guessing. The Viewer protocol remains `1.3.0`.
14
+
15
+ Project-aware `check` resolves and validates its immutable baseline before
16
+ candidate capture. Candidate URL, viewport, targets, and ordinary operational
17
+ settings still come from project configuration. The check replays only the
18
+ validated baseline request's optional `scrollScenario` and `explicitState`
19
+ through `normalizeRequest` and the canonical observer. Scroll is executed by
20
+ the normal browser capture path. `explicitState` replay retains the
21
+ caller-declared comparison identity only; it never infers or establishes
22
+ browser, application, theme, or session state. Project config remains schema
23
+ 1.1.0 and does not accept either field; `init` and normal `capture` remain
24
+ project-config-only.
25
+
26
+ The v0.10 Batch 4 actual-frontend entry route promotes only explicitly selected, saved, confirmed, canonically promotable runtime intent without activating project acceptance. It freezes the exact resulting contract instance with the saved annotation, source observation, and configured persistent baseline contract through the Batch 2 workflow-creation owner. Repeated equivalent starts share `contractRequestId` while receiving fresh contract, visual-change request, and workflow instance identities. Activation remains a later explicit workflow action.
27
+
28
+ Reference entry freezes only an exact approved reference, an annotation authored
29
+ against that exact instance, an exact baseline observation, and explicitly
30
+ authored valid bindings. Adequacy, compatibility, complete required-region
31
+ coverage, and canonical binding evaluation fail closed. Imported references,
32
+ inferred bindings, and session bindings are not executable workflow scope.
33
+
34
+ `VisualChangeAgentHandoff` uses handoff kind
35
+ `my-frontend-observer/visual-change-agent-handoff` and version `1.0.0`. It is a
36
+ bounded non-artifact transfer contract and is never discovered or persisted as
37
+ Observer evidence. It includes confirmed scope, Observer-owned bounded context,
38
+ optional unchanged supplemental context, the exact post-edit check instruction,
39
+ and optional non-authoritative orchestrator correlation.
40
+
41
+ The current cycle is derived from the latest attempt. A pending attempt blocks
42
+ another check or handoff until explicit correction, acceptance, or abandonment.
43
+ Only the latest canonical PASS may be accepted. Review changes only the latest
44
+ attempt's review object in a fresh workflow revision. Governance references may
45
+ be recorded only after acceptance and only for already-persisted canonical
46
+ approval results; they do not perform approval or change project configuration.
47
+ No existing evidence schema changed for v0.10.
48
+
49
+ ## Current contracts
50
+
51
+ The observation artifact contract is published in the current
52
+ `my-frontend-observer@0.7.0` package and proven both from the source checkout
53
+ and from the packed npm tarball, on Windows, Linux, and macOS. The observation
54
+ schema is `1.2.0` (see "v0.2 target contract" and "v0.3 scroll scenario
55
+ contract" below):
56
+
57
+ - artifact kind `my-frontend-observer/observation`, schema version `1.2.0`
58
+ (independent of the package version);
59
+ - one artifact root per observation, `<outputLocation>/<observationId>/`,
60
+ containing exactly `manifest.json` (the full `ObservationArtifact`, with
61
+ page/target evidence embedded inline) and `screenshot.png` - there is no
62
+ separate `evidence.json`;
63
+ - `manifest.json` is written last, after `screenshot.png`, via one atomic
64
+ directory rename, so a consumer never observes a partially-written
65
+ artifact; a filesystem failure anywhere in that sequence reports the
66
+ `artifact-write-failure` diagnostic and leaves no completed artifact;
67
+ - internal artifact references (e.g. `screenshot.png`) are relative to the
68
+ artifact root, never an absolute machine path; the observation's logical
69
+ identity is its `observationId`, not its filesystem location;
70
+ - evidence states `available`, `unavailable`, `not-applicable`, `partial`;
71
+ evidence sources `browser`, `computed-browser`, `derived`;
72
+ - a stable diagnostic vocabulary (`src/domain/diagnostics.ts`) and completion
73
+ states `complete`, `partial`, `warning`, `invalid-request`, `fatal`
74
+ (`src/domain/completion.ts`);
75
+ - observation/request identity, producer/package identity, and browser
76
+ provenance are all present in every persisted manifest.
77
+
78
+ This contract is implemented and published; no public programmatic-API
79
+ compatibility promise has been made for the observation engine itself. v0.6
80
+ additionally publishes the bounded-agent-context/correlation programmatic surface
81
+ described later in this document.
82
+
83
+ ## v0.2 target contract (shipped as part of this release)
84
+
85
+ v0.2 introduces a canonical target-configuration model: each configured target has a stable
86
+ observer-level `name` plus an ordered array of bounded `locators`
87
+ (`role`, `id`, `data-attribute`, `semantic-element`, `css`, `text`). This
88
+ identity is distinct from both the browser locator definition that resolves
89
+ it and any source-code identity. The legacy `{name, selector}` shape remains
90
+ accepted and normalizes to a one-item `css` locator, so every published
91
+ `0.1.0` CLI invocation continues to work unchanged. Locator precedence is the
92
+ configured array order; resolution stops on the first unique match, on any
93
+ ambiguous match (never falling through to a later locator), or on an
94
+ unevaluable locator - never silently. All six frozen locator kinds are now
95
+ resolved against a real Chromium page (`role` via Playwright's accessibility-
96
+ role/name locator with exact name matching, `id`/`data-attribute` via exact
97
+ CSS attribute-equals matching that never reinterprets the configured value as
98
+ selector syntax, `semantic-element` via the frozen tag set, `css` via the
99
+ existing v0.1 behavior, `text` via exact-text matching only); every kind
100
+ converges on the same measurement path, so locator strategy never changes the
101
+ resulting target evidence shape.
102
+
103
+ Each resolved target's evidence record additionally carries three bounded
104
+ fields: `semanticState` (a first family of `disabled`/`expanded`/
105
+ `checked`/`selected`/`pressed`/`current` values read from the element's own
106
+ native form-control properties and explicit `aria-*` attributes - a key is
107
+ present only when the browser exposes that state as applicable to this
108
+ element, so an explicit `false` is always distinguishable from "not
109
+ applicable"; `not-applicable` when no supported state applies at all);
110
+ `landmark` (derived only from the already-captured browser-exposed
111
+ role - never from locator kind or HTML tag - against the standard landmark
112
+ role set `banner`/`navigation`/`main`/`complementary`/`contentinfo`/`form`/
113
+ `region`/`search`); and `containment` (bounded DOM containment checked only
114
+ among the other explicitly configured targets in the same observation, in
115
+ configured order, never a layout/relationship graph - `available` when every
116
+ other configured target was itself resolved and checked, `partial` when one
117
+ or more could not be, `unavailable` when the target itself never resolved).
118
+ Stable observer target identity is proven, not just declared: the same
119
+ target configuration produces the same `requestId` across repeated
120
+ observations (with a fresh `observationId` each time); changing a target's
121
+ locator strategy while keeping its stable name changes `requestId` but not
122
+ the `targetEvidence` key; and actual runtime disappearance of a
123
+ still-configured target changes only its resolution status, never the
124
+ `requestId`.
125
+
126
+ The full canonical semantic target model above is reachable through the
127
+ real public CLI: `my-frontend-observer observe --targets-file <json-file>`
128
+ supplies the structured `{ "targets": [...] }` collection (see
129
+ `docs/COMMANDS.md` "Structured semantic targets") as an alternative to the
130
+ existing `--target id=css-selector` shorthand - the two are mutually
131
+ exclusive per invocation, and both converge on the same
132
+ `normalizeRequest()`/browser-resolver/artifact path, so a semantic
133
+ observation produces exactly the same `manifest.json` shape as a
134
+ CSS-shorthand one. Schema `1.1.0` was the v0.2 published artifact schema;
135
+ schema `1.2.0` has been emitted since v0.3 and remains the observation schema
136
+ in the current published v0.7.0 package, for both target-input modes
137
+ (target semantics are unchanged from v0.2 - see the v0.3 scroll scenario
138
+ contract below for what schema `1.2.0` actually adds). `--targets-file`'s
139
+ local input path is never part of the persisted request identity or
140
+ artifact.
141
+
142
+ ## v0.3 scroll scenario contract (shipped as part of this release)
143
+
144
+ v0.3 introduces one optional, additive request/evidence concern: a bounded
145
+ runtime scroll scenario, schema `1.2.0`.
146
+
147
+ A normalized request may carry `scrollScenario: { action }` with exactly one
148
+ of two frozen action kinds:
149
+
150
+ - `{ "kind": "window-scroll-by", "deltaX": <int>, "deltaY": <int> }`
151
+ - `{ "kind": "target-scroll-by", "target": "<stable target name>", "deltaX": <int>, "deltaY": <int> }`
152
+
153
+ `deltaX`/`deltaY` are signed integers bounded to `[-20000, 20000]`; at least
154
+ one must be non-zero. `target-scroll-by.target` refers only to an existing
155
+ stable configured target `name` (never a selector) and resolves through the
156
+ same canonical `resolveConfiguredTargets` algorithm every v0.2 locator kind
157
+ already uses - there is no second target-resolution path. A request with no
158
+ scenario normalizes and identifies exactly as it did before v0.3.
159
+
160
+ Execution (both action kinds share one code path): perform the immediate,
161
+ non-smooth scroll (`window.scrollBy`/`element.scrollBy`, `behavior:
162
+ 'instant'`) on the already-navigated, already-ready page; wait exactly two
163
+ `requestAnimationFrame` cycles; capture a final runtime snapshot. No second
164
+ browser, page, or navigation is ever created. The resulting scroll position
165
+ is browser-authoritative and may be clamped by document/element boundaries;
166
+ a scenario producing no movement is still a valid, successfully persisted
167
+ observation.
168
+
169
+ The scenario evidence lives entirely inside the existing `manifest.json` as
170
+ one additional optional `scrollScenarioEvidence` field on `ObservationArtifact`
171
+ - there is no separate `scroll.json`/`scenario.json`. It contains:
172
+
173
+ - `initial`/`final`: bounded `ScrollRuntimeSnapshot`s (window `scrollX`/
174
+ `scrollY`; the browser's own scrolling-root/`documentElement`/`body`
175
+ metrics; per-configured-target `scrollTop`/`scrollLeft`/`scrollWidth`/
176
+ `scrollHeight`/`clientWidth`/`clientHeight`, actual overflow, bounding
177
+ rectangle, and viewport relation);
178
+ - `transition`: bounded before/after change evidence (window scroll deltas;
179
+ per-target `scrollTop`/`scrollLeft`/bounding-position/viewport-relation
180
+ changes; `enteredViewport`/`leftViewport`) - never a generic recursive
181
+ diff, and a target is simply omitted when either side's evidence isn't
182
+ itself usable (e.g. it never resolved);
183
+ - `scrollOwner`: one derived `EvidenceField<ScrollOwnerInterpretation>`
184
+ (`document` | `target:<stable-name>` | `none` | `indeterminate`), always
185
+ `source: "derived"` with non-empty `derivedFrom` naming the exact
186
+ contributing scroll-position measurements. Ownership is derived only from
187
+ observed `scrollTop`/`scrollLeft`/`window.scrollX`/`window.scrollY`
188
+ changes - never from bounding-rectangle movement (which moves for every
189
+ configured target whenever the document scrolls), computed overflow,
190
+ `position: fixed`/`sticky`, or DOM hierarchy.
191
+
192
+ Actual dimensional overflow (`scrollWidth > clientWidth` /
193
+ `scrollHeight > clientHeight`) is always reported separately from the
194
+ computed `overflow-x`/`overflow-y` CSS declaration; a declared
195
+ `overflow: auto` container with content that fits produces
196
+ `horizontalOverflow`/`verticalOverflow: false`. Viewport relation
197
+ (`above`/`intersecting`/`below`, `intersectsViewport`, `fullyWithinViewport`)
198
+ is derived only from bounding geometry plus viewport size, relative to the
199
+ browser viewport; a hidden/non-rendered target's viewport relation is
200
+ `not-applicable`, never a fabricated geometry claim - hidden and offscreen
201
+ remain distinct evidence concepts, and the existing `target-hidden`
202
+ diagnostic is unaffected.
203
+
204
+ The ordinary, already-existing `pageEvidence`/`targetEvidence`/
205
+ `screenshot.png` for a scenario observation always describe this same final
206
+ post-action state, never the pre-action state.
207
+
208
+ The scenario request participates in `requestId`; the runtime result
209
+ (actual scroll distance, clamping, or scroll-owner outcome) never does. The
210
+ public entry point is `my-frontend-observer observe --scroll-scenario-file
211
+ <json-file>` (see `docs/COMMANDS.md`); the file supplies the scenario value
212
+ directly, and its local path is operational input only, exactly like
213
+ `--targets-file`'s path - never persisted, never part of request identity.
214
+
215
+ ## v0.4 comparison contract (shipped as part of this release)
216
+
217
+ **Current status: shipped as part of the published `my-frontend-observer@0.4.0`
218
+ package and unchanged through the current `0.7.0` release.** Observation
219
+ schema remains `1.2.0`. Comparison is a distinct artifact kind and schema,
220
+ never a bump to the observation schema:
221
+
222
+ - artifact kind: `my-frontend-observer/comparison`;
223
+ - comparison schema: `1.0.0`.
224
+
225
+ **Geometry tolerance**: `ComparisonConfig.geometryTolerancePx`, default
226
+ `0.5` CSS px, bounded `[0, 10]`. Suppresses insignificant subpixel noise
227
+ only - never a design contract, never permission for a change.
228
+
229
+ **Layout relationship graph**: `deriveLayoutRelationships(observation,
230
+ options?)` derives, per observation, a bounded `LayoutRelationshipGraph`
231
+ among configured targets only (≤20 targets, ≤190 unordered pairs):
232
+ horizontal order (`left-of`/`right-of`/`horizontally-overlapping`),
233
+ vertical order (`above`/`below`/`vertically-overlapping`), area overlap
234
+ (`overlaps`/`does-not-overlap`), relative width (`wider-than`/
235
+ `narrower-than`/`equal-width-within-tolerance`), geometric fit
236
+ (`fits-inside`/`does-not-fit-inside` - geometry-only, deliberately distinct
237
+ from DOM containment), vertical sequencing (`follows-vertically`), and one
238
+ page-level relationship (`document-width-fits-viewport`/
239
+ `document-width-exceeds-viewport`). Every relationship carries explicit
240
+ evidence-path provenance back to the source observation. A configured
241
+ target lacking usable geometry is listed as honestly unresolved
242
+ (`not-found`/`ambiguous`/`unavailable`/`hidden`), never fabricated as a
243
+ zero-sized region.
244
+
245
+ **Comparability**: evaluated before any rendered difference, using exactly
246
+ three states (`comparable`/`comparable-with-warnings`/`incomparable`) with
247
+ structured reasons, never a bare boolean. Hard incompatibilities (page URL,
248
+ viewport, browser engine, scroll-scenario configuration mismatch) force
249
+ `incomparable`; producer-version, browser-version, and target-configuration
250
+ differences are warning-only; theme/authenticated-state/application-state
251
+ identity are recorded as `unassessed` dimensions the observer does not yet
252
+ model - never silently claimed identical. An `incomparable` result still
253
+ persists a structurally valid `ComparisonArtifact` with empty rendered
254
+ differences, not a fabricated comparison.
255
+
256
+ **Difference categories**: `appeared`/`disappeared` (only for a stable
257
+ target name configured on both sides, transitioning between a definite
258
+ `not-found` and `matched` resolution status - never for a target merely
259
+ added/removed from configuration, which is its own separate
260
+ `configurationChanges` entry), `moved`/`resized` (tolerance-aware, a target
261
+ may be both), `visibility-changed`, `clipping-changed` (reusing the
262
+ canonical `deriveTargetClipping` helper, never re-derived), `horizontal-
263
+ overflow-changed`/`vertical-overflow-changed` (actual dimensional overflow,
264
+ reusing the existing `deriveOverflowEvidence` helper - never inferred from
265
+ a CSS declaration alone), `containment-changed` (reusing existing v0.2
266
+ `TargetContainment` evidence), `page-size-changed`, `scroll-owner-changed`
267
+ (comparing `scrollScenarioEvidence.scrollOwner` only when scenario
268
+ *configuration* already matched), `relative-position-changed` (a relation
269
+ in the horizontal-order/vertical-order/area-overlap families changed - kept
270
+ distinct from plain absolute target movement) and `relationship-changed`
271
+ (every other relationship-family transition). Relationship changes are
272
+ matched by structural identity (family + subject/related target, or the
273
+ page-level key), never by array position.
274
+
275
+ **Explicit dependency evidence**: `ComparisonConfig.expectedDependencies`
276
+ lets a caller declare an expected relationship between two targets' numeric
277
+ properties (`x`/`y`/`width`/`height`) and directions (`increase`/
278
+ `decrease`/`change`/`unchanged`), always carrying `source:
279
+ "explicit-config"`. The observer never synthesizes a declaration from
280
+ observed co-change. Each declaration evaluates independently to exactly one
281
+ of `consistent`/`not-observed`/`contradictory-to-declaration`/
282
+ `unavailable` - never a causal claim (no `causedBy`/`causalConfidence`/
283
+ `causalScore`/`dependencyStrength`) and never a PASS/FAIL/approval verdict.
284
+ That distinction (evidence vs. contract verdict) is the boundary between
285
+ v0.4 and v0.5+.
286
+
287
+ **Comparison identity**: `comparisonRequestId` is a pure, deterministic
288
+ function of `{beforeObservationId, afterObservationId, normalized
289
+ ComparisonConfig}` - direction-sensitive (`compare(A, B) !==
290
+ compare(B, A)`), and never includes an operational filesystem path.
291
+ `comparisonId` is fresh per execution (same pattern as `observationId`).
292
+
293
+ **Source references**: the comparison artifact retains enough logical
294
+ identity to trace back to its authoritative source observations
295
+ (`observationId`, `requestId`, `producer`, `observationSchemaVersion`, and
296
+ the source `screenshot.path`) without embedding the full
297
+ `ObservationArtifact` or copying screenshot bytes. The persisted comparison
298
+ directory contains `manifest.json` only.
299
+
300
+ The public entry point is `my-frontend-observer compare --before <root>
301
+ --after <root> --output <directory> [--config-file <json-file>]` (see
302
+ `docs/COMMANDS.md`) - comparison itself never launches a browser.
303
+
304
+ ## v0.5 frontend contract and evaluation (shipped as part of this release)
305
+
306
+ Downstream of the v0.4 observation/comparison/relationship evidence above,
307
+ `src/domain/frontendContracts.ts` freezes the v0.5 contract/change-scope
308
+ model, `src/domain/frontendContractIdentity.ts` freezes deterministic
309
+ contract/baseline/clause identity, and `src/domain/frontendContractEvaluation.ts`
310
+ implements the one canonical pure evaluation engine. Baseline/per-change
311
+ contract persistence, evaluation-artifact persistence, explicit baseline
312
+ approval, and public CLI exposure are all implemented and shipped (see
313
+ "v0.5 contract and evaluation persistence" and "v0.5 public contract/
314
+ evaluation commands" below).
315
+
316
+ **Contract classes**: a `PersistentBaselineContract` (append/supersession-based
317
+ history via an optional `supersedesBaselineId`) and a `PerChangeContract`
318
+ (the allowed scope of one requested change). Both share `artifactKind:
319
+ "my-frontend-observer/frontend-contract"` and `schemaVersion: "1.0.0"` - an
320
+ independent family from the observation (`1.2.0`) and comparison (`1.0.0`)
321
+ schemas; the frontend-contract schema constant happens to share the version
322
+ string `1.0.0` with comparison's by coincidence only.
323
+
324
+ **Four authored categories, one derived classification**: every per-change
325
+ clause is authored as exactly one of `requested`, `expected-dependent`,
326
+ `protected`, or `preserved`. `unexpected` is a fifth, *derived-only*
327
+ classification the evaluator produces for a meaningful rendered difference no
328
+ active clause accounts for - it can never be authored as a permission.
329
+
330
+ **Bounded contract primitives**: 15 frozen `ContractPrimitive` kinds cover
331
+ visibility, clipping, width bounds, non-overlap, relative width, vertical
332
+ sequence, geometric fit (explicitly distinct from DOM containment),
333
+ document-width-vs-viewport, scroll ownership, initial-viewport position,
334
+ relationship-unchanged, and property-unchanged/increases/decreases - a closed
335
+ vocabulary, never a generic expression language.
336
+
337
+ **Contract tolerance**: `exact` / `absolute-px` / `percent`, independent of
338
+ `ComparisonConfig.geometryTolerancePx` (which only suppresses insignificant
339
+ comparison noise and is never contract authorization). Percent tolerance's
340
+ denominator is the absolute before-value.
341
+
342
+ **Required vs. permitted expected-dependent**: `required` clauses must occur
343
+ compliantly to pass; `permitted` clauses accept no change or a compliant
344
+ change, and fail only on a strictly contradictory change.
345
+
346
+ **Evaluation result vocabulary**: each clause resolves to `pass` / `fail` /
347
+ `unavailable` (with a required non-empty reason - required evidence gaps and
348
+ an `incomparable` source comparison never fabricate a `pass`) / `conflict`
349
+ (with at least two `conflictingClauseIds` - covers both an unresolved
350
+ baseline/per-change contradiction and an unknown `supersedesBaselineClauseIds`
351
+ reference). The overall verdict is `PASS` only when every clause result is
352
+ `pass` and no unexpected change remains; otherwise `FAIL` - there is no
353
+ partial-pass scoring.
354
+
355
+ **Explicit supersession, never inferred**: a per-change clause may list
356
+ `supersedesBaselineClauseIds` to remove specific baseline clauses from active
357
+ evaluation. Two clauses that structurally contradict each other on the same
358
+ (target, property) without explicit supersession produce a `conflict`, never
359
+ a silent preference for one side.
360
+
361
+ **Reuses existing v0.4 evidence directly**: the evaluator consumes an
362
+ already-computed `ComparisonArtifact` (`differences`, `relationshipChanges`,
363
+ `relationshipsBefore`/`relationshipsAfter`, `comparability`) and the source
364
+ `ObservationArtifact` pair - it never re-launches a browser, re-resolves a
365
+ target, or reimplements clipping/relationship/scroll-owner derivation.
366
+ Unexpected-change derivation reads `ComparisonArtifact.differences` only
367
+ (which already includes one difference per relationship change), so a single
368
+ logical transition is never double-counted.
369
+
370
+ ## v0.5 contract and evaluation persistence (shipped as part of this release)
371
+
372
+ Persistence consumes the frozen v0.5 domain above; it never redefines it.
373
+ `src/artifacts/frontendContractArtifactWriter.ts`/`frontendContractArtifactReader.ts`
374
+ persist and read both `PersistentBaselineContract` and `PerChangeContract`
375
+ symmetrically (both already share `CONTRACT_ARTIFACT_KIND`/`CONTRACT_SCHEMA_VERSION`,
376
+ so one writer/reader pair serves both contract classes) as
377
+ `<outputLocation>/<baselineId|contractId>/manifest.json`, following the same
378
+ atomic-write discipline as `artifacts/artifactWriter.ts`/`artifacts/comparisonArtifactWriter.ts`
379
+ (sibling temporary directory, then one atomic rename; an existing directory at
380
+ the final identity is a genuine collision and is rejected, never overwritten -
381
+ prior baseline history is never rewritten). `src/artifacts/comparisonArtifactReader.ts`
382
+ is a new Batch 3 addition (no comparison reader existed before) mirroring
383
+ `artifacts/artifactReader.ts`'s discipline exactly, changing no comparison
384
+ semantics and keeping comparison schema `1.0.0`.
385
+
386
+ **Evaluation artifact envelope**: Batch 1 froze the evaluation-result
387
+ vocabulary (`ClauseEvaluationResult`, `OverallVerdict`) but not a persistable
388
+ envelope, so `src/domain/frontendContractEvaluationArtifact.ts` adds exactly
389
+ that - `artifactKind: "my-frontend-observer/frontend-contract-evaluation"`,
390
+ `schemaVersion: "1.0.0"` (its own independent family, distinct from
391
+ observation/comparison/frontend-contract), an `evaluationId`/`evaluationRequestId`
392
+ pair, bounded `before`/`after` source-observation references, and
393
+ `comparisonId`/`comparisonRequestId` plus `contracts: {baselineId,
394
+ contractId}` references - never an embedded `ObservationArtifact` or copied
395
+ screenshot. It reuses `ClauseEvaluationResult`/`OverallVerdict`/
396
+ `UnexpectedChangeResult` unchanged and contains no evaluation logic itself.
397
+ `evaluationRequestId` is a deterministic function of `{baselineId,
398
+ contractId, beforeObservationId, afterObservationId, comparisonRequestId}`
399
+ (`frontendContractIdentity.ts#buildFrontendContractEvaluationRequestIdentity` -
400
+ deliberately `comparisonRequestId`, not the fresh-per-execution
401
+ `comparisonId`, so semantically identical evaluations share an identity);
402
+ `evaluationId` reuses the existing generic `buildFrontendContractInstanceIdentity`
403
+ unchanged. `src/artifacts/frontendContractEvaluationArtifactWriter.ts`/
404
+ `frontendContractEvaluationArtifactReader.ts` persist/read it with the same
405
+ atomic-write discipline as above.
406
+
407
+ **Application seam**: `src/application/frontendContractEvaluationService.ts#evaluateAndPersist`
408
+ calls the existing pure `evaluateFrontendContract` exactly once and - only
409
+ for a structurally constructible result, whether the verdict is `PASS` or
410
+ `FAIL` - persists exactly one evaluation artifact; an `{ok: false}` evaluator
411
+ result (evidence could not be constructed into an evaluation at all) is never
412
+ persisted as a fabricated artifact. `evaluateAndPersistFromArtifactRoots` is
413
+ the future-CLI-facing wrapper: it reads two observations through the existing
414
+ `readObservationArtifact` (never a second observation reader), the
415
+ comparison and the two contracts through the readers above, then delegates
416
+ to `evaluateAndPersist` exactly once.
417
+
418
+ ## v0.5 public contract/evaluation commands (shipped as part of this release)
419
+
420
+ Three public commands expose the persistence/evaluation contract above (see
421
+ `docs/COMMANDS.md` for exact flags/output/exit behavior, not duplicated
422
+ here):
423
+
424
+ - `approve-baseline` → `frontendContractPersistenceService.ts#approveAndPersistBaseline`
425
+ → validates a `PersistentBaselineContract` and its `sourceObservation`
426
+ coherence against a supplied observation artifact → persists via
427
+ `frontendContractArtifactWriter.ts`. The only baseline-approval act in the
428
+ observer.
429
+ - `save-change-contract` → `frontendContractPersistenceService.ts#persistPerChangeContract`
430
+ → validates a `PerChangeContract` (rejecting a baseline contract, an
431
+ authored `unexpected` category, or any other structural violation) →
432
+ persists via the same writer. Persistence only, never approval.
433
+ - `evaluate-contract` → `frontendContractEvaluationService.ts#evaluateAndPersistFromArtifactRoots`
434
+ → `evaluateFrontendContract` exactly once → `frontendContractEvaluationArtifactWriter.ts`
435
+ exactly once. `--enforce` affects only the process exit status for an
436
+ already-persisted `FAIL` verdict.
437
+
438
+ No command infers baseline approval or supersession automatically - not
439
+ `compare`, not a `PASS` evaluation, not any artifact writer.
440
+
441
+ This full command sequence is proven against real Chromium observations (not
442
+ hand-constructed artifacts) - see "v0.5 real-browser workflow proof" below.
443
+
444
+ ## v0.5 real-browser workflow proof (shipped as part of this release)
445
+
446
+ `tests/browser/cliFrontendContracts.test.ts` and
447
+ `scripts/dev/builtCliFrontendContractsBrowserSmoke.mjs` drive the complete
448
+ `observe` → `approve-baseline` → `save-change-contract` → `observe` →
449
+ `compare` → `evaluate-contract` sequence against a real disposable local HTTP
450
+ fixture and real Chromium, proving two scenarios:
451
+
452
+ - a fully successful contract change - a real observed navigation-width
453
+ decrease and workspace-width increase, both satisfying their authored
454
+ `requested`/`expected-dependent` clauses, an unchanged `protected` rail
455
+ width, and an unclipped `preserved` navigation - overall `PASS`;
456
+ - the "milestone signature" failure - the same locally successful requested
457
+ change (navigation shrinks, workspace expands, both still `pass`)
458
+ co-occurring with a genuine `protected` right-rail width regression (a real
459
+ `resized` comparison difference) and a genuine `preserved` navigation
460
+ clipping regression (a real `clipping-changed` difference, `not-clipped` →
461
+ `clipped`) - overall `FAIL`.
462
+
463
+ Both scenarios confirm: `--enforce` changes only the process exit status
464
+ (`0` without it, nonzero with it) for the identical persisted
465
+ `evaluationRequestId`/`clauseResults`; every source observation and
466
+ comparison artifact is byte-identical before and after evaluation; the
467
+ evaluation directory contains `manifest.json` only (no copied screenshot);
468
+ and no operational filesystem path is ever serialized into a persisted
469
+ manifest. This is real-browser evidence layered on top of the CLI-level
470
+ proof in `tests/unit/cliFrontendContracts.test.ts` and the Chromium-free
471
+ `scripts/dev/builtCliFrontendContractsSmoke.mjs` - it does not replace them.
472
+
473
+ ## v0.6 bounded agent context and correlation contract (released as `0.6.0`)
474
+
475
+ **Current status: released as package version `0.6.0`, tag `v0.6.0`, from
476
+ the canonical `canonicalization/v0.6` lineage.** Bounded-agent-context is a new, independent
477
+ artifact-kind family, schema `1.0.0` (`BOUNDED_AGENT_CONTEXT_ARTIFACT_KIND =
478
+ "my-frontend-observer/bounded-agent-context"`) - never a bump to
479
+ observation/comparison/frontend-contract/evaluation schemas, which remain
480
+ `1.2.0`/`1.0.0`/`1.0.0`/`1.0.0` respectively. Unlike those families, there is
481
+ **no disk artifact writer/reader** for bounded-agent-context: it is a pure
482
+ programmatic contract and derivation layer, exported from `src/index.ts`
483
+ only.
484
+
485
+ **Bounded runtime projection**
486
+ (`src/domain/boundedAgentContextProjection.ts#projectBoundedAgentContext`)
487
+ produces a `BoundedRuntimeTargetProjection` from already-captured v0.1-v0.5
488
+ evidence, containing: page/viewport identity; stable target identities;
489
+ important geometry and runtime behavior; layout/behavior relationships;
490
+ before/after differences; contract clause results; requested/expected-
491
+ dependent/protected/preserved scope - reusing `src/domain/
492
+ frontendContracts.ts`'s existing clause types verbatim, never a
493
+ reimplementation; diagnostics; screenshot/artifact references; provenance;
494
+ and explicit `OmissionRecord`/`TruncationRecord` metadata with bounded
495
+ aggregate-cap summarization once a limit is reached.
496
+
497
+ **Adequacy**: every projection carries an `Adequacy` value
498
+ (`adequate`/`partial`/`inadequate`) plus a structured, closed
499
+ `ADEQUACY_REASON_CODES` vocabulary - evidence existing is not itself
500
+ adequacy; a required omission or an `incomparable`/unavailable upstream
501
+ source is reflected honestly rather than silently reported as sufficient.
502
+
503
+ **Runtime/static correlation**
504
+ (`src/domain/boundedAgentContextCorrelation.ts#deriveRuntimeStaticCorrelations`/
505
+ `attachRuntimeStaticCorrelations`) evaluates each stable runtime target
506
+ against caller-supplied candidate static-evidence records into exactly one of
507
+ three outcomes: `correlated`, `ambiguous` (multiple competing candidates,
508
+ all preserved and visible - never silently resolved to one), or
509
+ `unavailable` (no supported candidate). The module accepts only plain,
510
+ already-retrieved candidate records and has no dependency on
511
+ `@dailephd/my-dev-kit` - the audit preceding implementation found no generic
512
+ static-side retrieval capability actually missing (see `docs/ROADMAP.md` v0.6
513
+ "Dependency direction"). A runtime target identity is carried through
514
+ verbatim; correlation never produces a `sourceOwner`/`causedBy`-shaped field,
515
+ so a stable runtime identity is never silently reported as source ownership.
516
+
517
+ **Identity**: `src/domain/boundedAgentContextIdentity.ts#buildBoundedAgentContextRequestIdentity`/
518
+ `buildBoundedAgentContextInstanceIdentity` follow the same
519
+ canonicalize+sha256(+opaque-nonce) pattern as `comparisonIdentity.ts`/
520
+ `frontendContractIdentity.ts`: a deterministic logical identity distinct from
521
+ a fresh per-execution instance identity.
522
+
523
+ **Export/public boundary**: `src/index.ts` exports the complete
524
+ bounded-agent-context/correlation type and function surface as a
525
+ programmatic library contract. There is no CLI command (`observe`/`compare`/
526
+ `approve-baseline`/`save-change-contract`/`evaluate-contract` remain the only
527
+ public commands) and no orchestrator/lab code in this repository - bounded
528
+ runtime-evidence consumption by `my-dev-kit-orchestrator` and exact
529
+ readers/fixtures/evaluation in `my-dev-kit-lab` are separate sibling-
530
+ repository deliverables outside `my-frontend-observer`'s public surface.
531
+
532
+ **Compatibility evidence**: cross-repository neutral verification (observer
533
+ `514bf3bb513764815a0a5b9e508d5836aa7d7fd8`, orchestrator `9473e4c`, lab
534
+ `271e72c`) passed with 6/6 requirement coverage and no known product
535
+ blockers; on the canonical worktree, `npm run typecheck`, `npm run lint`,
536
+ `npm test` (627 tests), `npm run test:browser` (120 tests), `npm run
537
+ test:security`, `npm run build`, and `npm run check:docs` all pass.
538
+
539
+ ## v0.7 external visual-reference contract direction (preserved through released v0.10)
540
+
541
+ External visual-reference support is released as package version `0.7.0`
542
+ (see "v0.7 Prompt 1" through "v0.7 Prompt 8" below for the exact contract).
543
+ The exact public type names, artifact kinds, schema versions, persistence
544
+ layout, and command/programmatic entry points were designed during v0.7
545
+ implementation from current repository precedent, following the constraints
546
+ below. v0.8 (released as package version `0.8.0` - see
547
+ `docs/CURRENT_STATE.md`) has preserved them. v0.9 is released and preserves
548
+ them. The implemented and released v0.10 workflow preserves them too.
549
+
550
+ **Distinct evidence domain**: an external reference is desired-design evidence,
551
+ not an `ObservationArtifact` and not the "before" side of a v0.4
552
+ `ComparisonArtifact`. Reference design vs candidate is distinct from both
553
+ before vs after comparison and frontend-contract evaluation. The implementation
554
+ must not fake this distinction by wrapping a raster image in an observation
555
+ shape.
556
+
557
+ **Reference identity and provenance**: a future reference contract must preserve
558
+ a deterministic logical reference identity/version where appropriate, source
559
+ image reference plus dimensions/format, provenance, bounded region definitions,
560
+ applicable viewport/theme/application-state identity, authored design intent,
561
+ relationship/style evidence where supported, limits/diagnostics, and approval/
562
+ supersession history. Operational filesystem paths must not become semantic
563
+ identity. A raw imported image never silently becomes an approved active
564
+ reference.
565
+
566
+ **Reference regions and runtime targets stay distinct**: a reference region
567
+ must have its own identity and coordinate semantics. Reference-region to runtime-
568
+ target association must be explicit and capable of representing ambiguity or
569
+ unavailability. Runtime target identity and reference identity must never
570
+ silently become static source ownership; static association still goes through
571
+ the v0.6 runtime/static correlation boundary.
572
+
573
+ **Applicability before fidelity**: viewport, theme, application state, and
574
+ other selected compatibility dimensions must be evaluated before ordinary
575
+ reference/candidate differences are produced. If the reference and candidate
576
+ represent different intended states, the result must be explicitly incompatible
577
+ or incomparable rather than filled with fabricated visual failures. Planning
578
+ should reuse or extend the canonical v0.4 comparability conventions where they
579
+ mean the same thing rather than invent an unrelated reference-only state model.
580
+
581
+ **Canonical contract semantics remain authoritative**: executable reference-
582
+ derived requirements must map into the existing v0.5 authored categories
583
+ `requested`, `expected-dependent`, `protected`, or `preserved`. The derived-only
584
+ `unexpected` classification remains derived-only. Informational or unassessed
585
+ reference evidence may stay outside executable contract evaluation until
586
+ explicitly promoted. A second reference-only PASS/FAIL taxonomy is forbidden.
587
+
588
+ **Tolerance separation**: reference-fidelity tolerances are not automatically
589
+ the same as v0.4 `ComparisonConfig.geometryTolerancePx` or v0.5 contract
590
+ tolerances. Planning must define property-specific semantics for reference
591
+ geometry, spacing, selected style evidence, text/font rendering differences,
592
+ asset-sensitive regions, and optional image similarity. One global pixel-perfect
593
+ threshold is not an acceptable contract.
594
+
595
+ **Structured evidence first**: geometry, relationships, authored requirements,
596
+ applicability, provenance, and selected bounded style/asset evidence remain
597
+ inspectable primary evidence. Screenshot-region or image-similarity evidence may
598
+ supplement them where reliable, but pixel similarity alone must not determine
599
+ success and must never override active baseline/per-change contracts.
600
+
601
+ **Bounded correction evidence**: future reference/candidate results must support
602
+ a bounded projection suitable for coding-agent correction, such as reference
603
+ measurement, candidate measurement, delta, failed relationship/style condition,
604
+ relevant reference/runtime identities, provenance, and active protected/
605
+ preserved constraints. Heavy reference image bytes should be referenced, not
606
+ copied into every downstream context packet.
607
+
608
+ **Approval and supersession**: reference import, reference approval, baseline
609
+ approval, reference supersession, and baseline supersession are separate acts.
610
+ A reference-fidelity `PASS`, a frontend-contract `PASS`, or a successful
611
+ before/after comparison must not silently approve or replace any reference or
612
+ baseline.
613
+
614
+ The v0.8 viewer, released as package version `0.8.0`, consumes this v0.7
615
+ reference/evaluation contract exactly as required - it creates no UI-only
616
+ reference model (see `docs/ARCHITECTURE.md` "v0.8 Batch 5"/"v0.8 Batch 6"
617
+ and `docs/reports/v0.8-reference-candidate-inspection-batch5.md`). v0.9
618
+ annotations (released in `0.9.0`) originate from runtime
619
+ screenshots or external references, preserve which source identity and
620
+ coordinate system they belong to, and feed the same canonical contract and
621
+ reference semantics - see "v0.9 visual annotation contract" below. v0.10
622
+ combines both entry modes into the full correction/approval workflow.
623
+
624
+ ## v0.9 visual annotation contract (released in 0.9.0)
625
+
626
+ v0.9 is released as `@dailephd/my-frontend-observer@0.9.0`.
627
+
628
+ **Artifact**: `VisualAnnotationArtifact`, artifact kind
629
+ `my-frontend-observer/visual-annotation`, schema version `1.0.0`. It stores one
630
+ exact canonical source (a runtime observation, or an imported or approved
631
+ external reference), the source coordinate space (runtime CSS pixels or
632
+ reference-image pixels), and bounded structured items. Each item has a stable
633
+ `annotationItemId`, one mark (point, rectangle, line, arrow, or note), an
634
+ optional explicit association, and an interpretation. A revision sets
635
+ `supersedesAnnotationId` and never rewrites its parent. The overlay SVG is
636
+ derived from the artifact and verified before it is served.
637
+
638
+ **An annotation is not a contract.** Saving an annotation never creates a
639
+ contract clause, a reference requirement, or a PASS/FAIL rule. Marks and
640
+ visible pixels are evidence. Overlap between a mark and a target or region is
641
+ not ownership and never creates an association.
642
+
643
+ **Interpretation states**: `uninterpreted`, `candidate`, and `confirmed`.
644
+ Confirmation is explicit and records `confirmedAt`. Editing the mark,
645
+ association, or intent of a confirmed item withdraws the confirmation.
646
+
647
+ **Runtime intent to contract**:
648
+
649
+ - Only selected, confirmed, supported runtime intent is promoted. Promotion
650
+ creates one normal canonical `PerChangeContract` through the existing
651
+ contract persistence service.
652
+ - Supported mappings use the existing `ContractPrimitive` vocabulary only.
653
+ `move` maps to `property-increases`/`property-decreases` on `x` or `y`.
654
+ `resize` maps to `property-increases`/`property-decreases` on `width` or
655
+ `height`. `preserve` maps to `property-unchanged-within-tolerance` for a
656
+ target property, or to `relationship-unchanged` for an explicitly associated
657
+ canonical relationship.
658
+ - Categories are the canonical `requested`, `expected-dependent` (with a
659
+ required `required` or `permitted` mode), `protected`, and `preserved`.
660
+ `unexpected` is never authored; it stays evaluator-derived.
661
+ - `remove` can be confirmed and saved, but it is not promotable in v0.9. The
662
+ contract vocabulary has no target-absent primitive, and no approximate
663
+ clause is fabricated.
664
+ - `inspect` intent and notes are informational and never promoted.
665
+ - Each clause's `supportingEvidence` records the annotation source and item
666
+ paths. Promotion never activates the contract unless explicitly requested,
667
+ and it never approves a baseline.
668
+
669
+ **Reference intent to a new reference revision**:
670
+
671
+ - Only selected, confirmed `reference-region` (`create` or `refine`) and
672
+ `reference-requirement` items are materialized. `inspect` and
673
+ `asset-sensitive` intent stay informational.
674
+ - The result is a new imported `ExternalReferenceArtifact`, created through the
675
+ existing canonical import service. It supersedes the selected source
676
+ reference (the approved reference id when the source is approved) and has
677
+ lifecycle `imported`. It is never approved automatically.
678
+ - Source regions keep their order. A refine replaces the rectangle of an
679
+ existing source region and keeps its id. Creates are appended in selection
680
+ order. Duplicate creates, duplicate refines, and create-then-refine in one
681
+ request are rejected.
682
+ - Source requirements stay first with identical recomputed ids. Selected
683
+ requirements are appended in selection order and validated against the final
684
+ regions. A selected relationship requirement must still hold for the final
685
+ geometry. Measurement requirements store only the subject and tolerance.
686
+ - Applicability, label, and the exact source image bytes are preserved. The
687
+ source reference is never modified, and project reference acceptance is not
688
+ changed.
689
+
690
+ Existing contract, reference relationship, adequacy, and fidelity evaluators
691
+ remain the only source of verdicts. v0.9 adds no annotation evaluator.
692
+
693
+ ## Approved v0.1 design inputs
694
+
695
+ The historical greenfield scaffold plan recorded these v0.1 design decisions:
696
+
697
+ - artifact kind `my-frontend-observer/observation`;
698
+ - schema version `1.0.0`, independent of package version;
699
+ - one portable directory containing `manifest.json`, `evidence.json`, and
700
+ `screenshot.png`;
701
+ - evidence states `available`, `unavailable`, `not-applicable`, and `partial`;
702
+ - evidence sources `browser`, `computed-browser`, and `derived`;
703
+ - bounded explicitly requested targets, provenance, diagnostics, completion
704
+ state, limits, and relative artifact references.
705
+
706
+ These were planning inputs only at the time they were recorded. As shown in
707
+ "Current contracts" above, the implemented contract matches them except for
708
+ the file layout: there is no separate `evidence.json` - page/target evidence
709
+ is embedded directly inside `manifest.json`.
710
+
711
+ Comparison and relationship contracts belong to v0.4, and canonical
712
+ change-scope contracts belong to v0.5 - see "v0.5 frontend contract and
713
+ evaluation" above for the full shipped contract model, identity, evaluation
714
+ engine, persistence, baseline approval, and CLI exposure. Bounded
715
+ agent-context and runtime/static correlation contracts are v0.6 - see "v0.6
716
+ bounded agent context and correlation contract" above for the full released
717
+ model. The text/config-driven coding-agent review plus non-graphical external
718
+ visual-reference foundation is v0.7 - see "v0.7 Prompt 1 external-reference
719
+ artifact contract" below for the foundation layer implemented so far. Viewer
720
+ consumption of that reference model is released in v0.8 (see "v0.7 external
721
+ visual-reference contract direction" above); dual-context annotation follows
722
+ in v0.9; both visual entry modes converge with the existing workflow in
723
+ v0.10.
724
+
725
+ ## v0.7 Prompt 1 external-reference artifact contract
726
+
727
+ Released as `0.7.0`. This is the foundation layer only: identity,
728
+ provenance, bounded image metadata, and a two-state lifecycle for one
729
+ externally supplied design-reference image. It implements no region,
730
+ geometry, relationship, requirement, tolerance, binding, or fidelity-
731
+ evaluation contract - those belong to later v0.7 prompts.
732
+
733
+ An external reference is a distinct evidence root, not a variant of
734
+ `ObservationArtifact`: it never reuses `ARTIFACT_KIND`/`SCHEMA_VERSION`
735
+ (observation), `COMPARISON_ARTIFACT_KIND`, or `CONTRACT_ARTIFACT_KIND`, and
736
+ those existing types gain no new field from this contract.
737
+
738
+ ```ts
739
+ const EXTERNAL_REFERENCE_ARTIFACT_KIND = 'my-frontend-observer/external-reference';
740
+ const EXTERNAL_REFERENCE_SCHEMA_VERSION = '1.0.0'; // independent of package.json version and every other family's schema version
741
+
742
+ type ExternalReferenceImageFormat = 'png' | 'jpeg' | 'webp';
743
+
744
+ interface ExternalReferenceImageReference {
745
+ path: string; // bare relative filename within the artifact's own directory
746
+ format: ExternalReferenceImageFormat;
747
+ width: number;
748
+ height: number;
749
+ byteLength: number;
750
+ sha256: string; // identity-bearing content hash of the raw image bytes
751
+ }
752
+
753
+ // Points back to the imported artifact that owns the image, without copying its bytes - mirrors ComparisonSourceObservationReference.
754
+ interface ExternalReferenceSourceReference {
755
+ referenceId: string;
756
+ referenceRequestId: string;
757
+ producer: { name: 'my-frontend-observer'; version: string };
758
+ schemaVersion: '1.0.0';
759
+ image: ExternalReferenceImageReference;
760
+ }
761
+
762
+ // Exactly two persisted states - no literal 'superseded' variant (see below).
763
+ type ExternalReferenceLifecycleState = { state: 'imported' } | { state: 'approved'; approvedAt: string };
764
+
765
+ interface ExternalReferenceArtifactBase {
766
+ artifactKind: 'my-frontend-observer/external-reference';
767
+ schemaVersion: '1.0.0';
768
+ referenceRequestId: string; // deterministic logical identity - shared by an imported artifact and every artifact produced by approving it
769
+ referenceId: string; // fresh per-persisted-instance identity
770
+ producer: { name: 'my-frontend-observer'; version: string };
771
+ provenance: { importedAt: string; label?: string };
772
+ supersedesReferenceId?: string; // explicit, forward-only supersession of a prior reference's referenceId
773
+ diagnostics: Diagnostic[];
774
+ completion: CompletionState;
775
+ }
776
+
777
+ // lifecycle.state === 'imported': owns the image.
778
+ interface ImportedExternalReferenceArtifact extends ExternalReferenceArtifactBase {
779
+ lifecycle: { state: 'imported' };
780
+ image: ExternalReferenceImageReference;
781
+ }
782
+
783
+ // lifecycle.state === 'approved': references, never copies, the imported artifact's image.
784
+ interface ApprovedExternalReferenceArtifact extends ExternalReferenceArtifactBase {
785
+ lifecycle: { state: 'approved'; approvedAt: string };
786
+ sourceReference: ExternalReferenceSourceReference;
787
+ }
788
+
789
+ type ExternalReferenceArtifact = ImportedExternalReferenceArtifact | ApprovedExternalReferenceArtifact;
790
+ ```
791
+
792
+ Key rules:
793
+
794
+ - `referenceRequestId` is a pure function of `{imageSha256, format, width,
795
+ height, supersedesReferenceId}` only - never a filesystem path, output
796
+ location, label, or timestamp. Byte-identical image content imported from a
797
+ different operational root produces the same `referenceRequestId`;
798
+ changing any of those fields changes it.
799
+ - `referenceId` is fresh (nonce-based) on every persisted write, including
800
+ every approval of an already-imported reference.
801
+ - Importing an image never approves it (`lifecycle.state` is always
802
+ `'imported'` immediately after import, regardless of a supplied label or
803
+ supersession target). Approval is a single explicit act
804
+ (`approveExternalReference`, mirroring `approveAndPersistBaseline`) that
805
+ refuses anything not currently in the `'imported'` state.
806
+ - Approving persists a *new* artifact instance (same `referenceRequestId`,
807
+ fresh `referenceId`) carrying a `sourceReference` back to the imported
808
+ artifact - it never mutates the imported artifact's own manifest, and never
809
+ copies the image bytes a second time.
810
+ - Supersession is represented only as a forward pointer
811
+ (`supersedesReferenceId` on the newer artifact); there is deliberately no
812
+ literal `'superseded'` lifecycle state, so an existing persisted artifact's
813
+ own manifest is never rewritten - immutability holds unconditionally rather
814
+ than depending on careful mutation discipline.
815
+ - Supported formats are frozen to exactly `png`/`jpeg`/`webp`, detected from
816
+ header/magic bytes only (never a caller-declared file extension), bounded
817
+ to `EXTERNAL_REFERENCE_MAX_IMAGE_BYTES` (20,000,000 bytes) and
818
+ `[EXTERNAL_REFERENCE_MIN_DIMENSION_PX, EXTERNAL_REFERENCE_MAX_DIMENSION_PX]`
819
+ (`[1, 8192]`) pixels per side. No OCR, no raster decode, no computer
820
+ vision, no automatic region detection.
821
+
822
+ Persisted as `<outputLocation>/<referenceId>/manifest.json` (+
823
+ `reference.<ext>` for an `'imported'` artifact only), via the same atomic
824
+ temp-dir-then-rename discipline as every other artifact family
825
+ (`src/artifacts/externalReferenceArtifactWriter.ts` /
826
+ `externalReferenceArtifactReader.ts`). CLI: `import-reference <image-file>
827
+ --output <dir> [--label] [--supersedes <root>]` and `approve-reference
828
+ --reference <root> --output <dir> [--supersedes <root>]`.
829
+
830
+ ## v0.7 Prompt 2 explicit reference regions and relationships
831
+
832
+ Released as `0.7.0`. Additive extension of the Prompt 1 contract above:
833
+ one new, optional `regions?: ReferenceRegion[]` field on
834
+ `ExternalReferenceArtifact` (both lifecycle variants), plus a pure,
835
+ non-persisted relationship-derivation capability. No schema version bump -
836
+ `EXTERNAL_REFERENCE_SCHEMA_VERSION` remains `'1.0.0'`, because the field is
837
+ genuinely optional/additive and every Prompt 1 artifact (which predates this
838
+ field entirely) remains valid without it.
839
+
840
+ ```ts
841
+ // domain/externalReferenceRegions.ts
842
+ interface ReferenceRegionRectangle { x: number; y: number; width: number; height: number; }
843
+ interface ReferenceRegion { id: string; rectangle: ReferenceRegionRectangle; }
844
+
845
+ // Pure derived geometry - never persisted, always recomputed, so it can never drift from the rectangle above.
846
+ interface ReferenceRegionGeometry {
847
+ x: number; y: number; width: number; height: number;
848
+ right: number; bottom: number; centerX: number; centerY: number;
849
+ }
850
+
851
+ const REFERENCE_REGION_ID_PATTERN = /^[A-Za-z0-9_-]{1,64}$/; // same convention as request/request.ts's target-name pattern
852
+ const MAX_REFERENCE_REGIONS = 20; // same bound value as request/request.ts's MAX_TARGETS - independently owned, coincidentally equal
853
+ ```
854
+
855
+ Region coordinate semantics: origin at the reference image's top-left
856
+ corner, x increasing rightward, y increasing downward, unit is
857
+ reference-image pixels (explicitly not CSS pixels - a static image has no
858
+ CSS box model), coordinates may be fractional. A region's rectangle must lie
859
+ entirely within its owning image's own already-validated
860
+ width/height - out-of-bounds geometry is rejected outright, never clamped.
861
+
862
+ Key rules:
863
+
864
+ - Only `{x, y, width, height}` is canonical/authored. `right`, `bottom`,
865
+ `centerX`, `centerY` are pure calculations over it
866
+ (`deriveReferenceRegionGeometry`) - never a second, potentially-drifting
867
+ stored copy of the same fact.
868
+ - Region content is identity-bearing:
869
+ `buildExternalReferenceRequestIdentity` gained an additive, optional
870
+ trailing `regions` parameter. Omitting it entirely (every Prompt 1 call
871
+ site, and any Prompt 2 call that legitimately has no regions) produces the
872
+ byte-identical hash Prompt 1 already produced - the parameter is left out
873
+ of the hashed view rather than defaulted to `null`, unlike
874
+ `supersedesReferenceId`. Authored region order participates in identity
875
+ (arrays are never reordered by the shared `canonicalize()`), mirroring
876
+ `domain/identity.ts`'s treatment of configured targets.
877
+ - Region IDs are unique case-insensitively within one artifact (mirroring
878
+ `request/request.ts`'s target-name dedup convention exactly).
879
+ - `import-reference` gained an optional `--regions-file <json-file>` of the
880
+ form `{ "regions": [...] }` (same object-root-wrapper convention as
881
+ `--targets-file`); a legacy invocation without it behaves exactly as in
882
+ Prompt 1. `approve-reference` carries an imported artifact's `regions`
883
+ forward verbatim (never re-validated, never re-derived, never dropped) -
884
+ approval never adds, removes, or edits regions.
885
+ - One new diagnostic code, `invalid-reference-region` (error), covers every
886
+ region-validation failure (missing/duplicate/malformed id,
887
+ non-finite/negative/zero geometry, out-of-image-bounds, over the bounded
888
+ region count) - deliberately not split into several codes, per the
889
+ "don't proliferate diagnostics" convention.
890
+
891
+ Reference-region relationships (`domain/externalReferenceRegionRelationships.ts`)
892
+ reuse the exact same pure, tolerance-aware geometry predicates that
893
+ `domain/relationships.ts#deriveLayoutRelationships` uses for runtime targets
894
+ (`horizontalOrderOf`/`verticalOrderOf`/`areaOverlapOf`/`relativeWidthOf`/
895
+ `geometricFitOf`/`verticalSequenceOf`, now exported additively from that
896
+ module with unchanged formulas) and the same `PairwiseRelationshipKind`
897
+ vocabulary and `EvidenceReference` type - never a duplicated or
898
+ reinterpreted copy. Only the six geometry-only families apply (horizontal
899
+ order, vertical order, area overlap, relative width, geometric fit, vertical
900
+ sequencing); DOM containment, scroll ownership, runtime visibility, and
901
+ page-width-vs-viewport are runtime/browser concepts with no reference-image
902
+ equivalent and are not reused. `fits-inside`/`does-not-fit-inside` is
903
+ geometry-only fit - it never claims DOM containment, which an external image
904
+ cannot expose.
905
+
906
+ ```ts
907
+ interface ReferenceRegionRelationship {
908
+ kind: PairwiseRelationshipKind;
909
+ subjectRegion: string; // deliberately distinct field name from PairwiseLayoutRelationship's subjectTarget
910
+ relatedRegion: string;
911
+ evidence: EvidenceReference[]; // e.g. { path: 'regions.header.rectangle' } - never a targetEvidence/browser path
912
+ }
913
+ ```
914
+
915
+ Relationships are **not persisted** on the artifact - `deriveReferenceRegionRelationships(referenceRequestId, regions, options)`
916
+ is a pure, deterministic, synchronous function any caller (a future prompt,
917
+ a test) calls on demand against an artifact's own `regions` field, avoiding
918
+ any possibility of a persisted relationship graph drifting from the region
919
+ data it was derived from. Bounded at `MAX_REFERENCE_REGIONS` regions ->
920
+ `MAX_REFERENCE_REGION_PAIRS` pairs `x` 6 families =
921
+ `MAX_REFERENCE_REGION_RELATIONSHIP_RECORDS` records maximum - the same
922
+ bounding shape as `relationships.ts`'s `MAX_PAIRWISE_RELATIONSHIP_RECORDS`.
923
+ This is a maximum capacity, never a required minimum region count - there is
924
+ no contract requiring any specific number of authored regions.
925
+
926
+ A reference relationship is a fact about the reference image's geometry
927
+ only. It is not a design requirement, not a pass/fail verdict, and does not
928
+ claim a runtime target or source owner exists - see
929
+ `docs/WORKFLOWS.md` "Current external-reference foundation workflow" for
930
+ where those later concepts (Prompt 3+) will attach.
931
+
932
+ ## v0.7 Prompt 3 selected design requirements, tolerance semantics, and reference-evidence adequacy
933
+
934
+ Released as `0.7.0`. Additive extension of the Prompt 1/2 contracts
935
+ above: one new, optional `requirements?: ExternalReferenceRequirement[]`
936
+ field on `ExternalReferenceArtifact` (both lifecycle variants). No schema
937
+ version bump - same reasoning as Prompt 2's `regions` field.
938
+
939
+ **Central distinction**: a region's geometry is REFERENCE EVIDENCE -
940
+ everything visibly/measurably present in the image. A requirement is
941
+ SELECTED DESIGN INTENT - only what the user/configuration explicitly chose
942
+ as mattering for later candidate evaluation. Nothing in this repository ever
943
+ turns a region property or a derived relationship into a requirement
944
+ automatically.
945
+
946
+ ```ts
947
+ // domain/externalReferenceRequirements.ts
948
+
949
+ // Reused directly from domain/frontendContracts.ts - not reinvented as a
950
+ // "reference-only" taxonomy; that type carries no runtime-only coupling.
951
+ // 'unexpected' remains impossible to author (not a member of this union).
952
+ type AuthoredChangeScopeCategory = 'requested' | 'expected-dependent' | 'protected' | 'preserved';
953
+ type ExpectedDependentMode = 'required' | 'permitted'; // required only (and exactly) when category === 'expected-dependent'
954
+
955
+ type ReferenceRequirementRegionProperty = 'x' | 'y' | 'width' | 'height' | 'right' | 'bottom' | 'centerX' | 'centerY'; // exactly ReferenceRegionGeometry's own fields
956
+ type ReferenceRequirementMeasurement = 'vertical-gap' | 'horizontal-gap' | 'center-x-delta' | 'center-y-delta' | 'left-edge-delta' | 'right-edge-delta';
957
+
958
+ type ReferenceRequirementSubject =
959
+ | { kind: 'region-property'; region: string; property: ReferenceRequirementRegionProperty }
960
+ | { kind: 'region-relationship'; subjectRegion: string; relatedRegion: string; relationship: PairwiseRelationshipKind } // reused from relationships.ts - geometry-only families only
961
+ | { kind: 'region-measurement'; subjectRegion: string; relatedRegion: string; measurement: ReferenceRequirementMeasurement };
962
+
963
+ // Deliberately NOT a reuse of frontendContracts.ts's ContractTolerance: that
964
+ // type's 'absolute-px' is implicitly runtime/CSS pixels. Reference-image
965
+ // pixels are a distinct, explicitly-labeled unit - nothing here assumes
966
+ // 1 reference pixel = 1 CSS pixel (Prompt 6 will need an explicit mapping).
967
+ type ReferenceRequirementTolerance = { kind: 'exact' } | { kind: 'absolute-reference-px'; amount: number } | { kind: 'percent'; amount: number };
968
+
969
+ interface ExternalReferenceRequirement {
970
+ requirementId: string; // system-computed from {subject, category, expectedDependentMode, tolerance} only - never authored
971
+ category: AuthoredChangeScopeCategory;
972
+ expectedDependentMode?: ExpectedDependentMode;
973
+ subject: ReferenceRequirementSubject;
974
+ tolerance?: ReferenceRequirementTolerance; // required for region-property/region-measurement; must be absent for region-relationship
975
+ }
976
+ ```
977
+
978
+ Key rules:
979
+
980
+ - Requirement identity (`requirementId`) is always system-computed
981
+ (`buildReferenceRequirementIdentity`, mirroring
982
+ `frontendContractIdentity.ts#buildClauseIdentity`'s exact shape) - the raw
983
+ authored input (`RawReferenceRequirement`) has no `requirementId` field at
984
+ all, and supplying one is a validation error. Unlike v0.5's
985
+ `BaselineClause`/`PerChangeClause` (which need an author-visible `clauseId`
986
+ for cross-document `supersedesBaselineClauseIds` references), Prompt 3
987
+ requirements have no cross-document reference need yet, so trusting an
988
+ authored id would only invite drift between a user-typed id and the
989
+ content it claims to identify.
990
+ - The reference-side expected value/relationship is never stored on the
991
+ requirement or the artifact - `deriveReferenceRequirementExpectation()` is
992
+ a pure function computed on demand from the artifact's own `regions`,
993
+ eliminating the exact drift risk of persisting e.g. `width: 424` alongside
994
+ a region whose rectangle could (in principle) later disagree with it.
995
+ - A requirement referencing a region id that does not exist in the
996
+ artifact's own `regions` is a **structural validation failure** (rejected
997
+ at construction/import time), never merely "unavailable" reference
998
+ evidence - `isValidReferenceRequirements` checks this before any
999
+ requirement reaches adequacy computation.
1000
+ - **Duplicate/conflicting subject rule**: no two requirements in one
1001
+ collection may share the same structural subject (same region+property,
1002
+ or the same unordered region pair + relationship, or + measurement),
1003
+ regardless of category. This single rule covers both "duplicate
1004
+ requirement" and "conflicting categories on the same subject" (e.g. the
1005
+ same region/property authored as both `requested` and `protected`) -
1006
+ v0.5's `evaluateFrontendContract#primitivesConflict` is a *runtime-
1007
+ evaluation-time* detector (it needs before/after `ObservationArtifact`
1008
+ evidence that does not exist yet at this stage) and could not be reused
1009
+ safely; Prompt 3 restricts invalid combinations at authoring time instead,
1010
+ per the documented precedent-review outcome.
1011
+ - Bounded at `MAX_REFERENCE_REQUIREMENTS` (50) requirements per artifact -
1012
+ a maximum capacity, never a required minimum (there is no contract
1013
+ requiring any specific number of authored requirements).
1014
+ - One new diagnostic code, `invalid-reference-requirement` (error), covers
1015
+ every requirement-authoring validation failure - deliberately not split
1016
+ further, per the "don't proliferate diagnostics" convention already used
1017
+ for `invalid-reference-region`.
1018
+ - `import-reference` gained an optional `--requirements-file <json-file>`
1019
+ (`{ "requirements": [...] }`, same object-root-wrapper convention as
1020
+ `--regions-file`/`--targets-file`); `approve-reference` carries an
1021
+ imported artifact's `requirements` forward verbatim (never re-validated,
1022
+ never re-derived, never dropped), exactly mirroring how it already
1023
+ handles `regions`.
1024
+
1025
+ **Reference-evidence adequacy** (`deriveReferenceRequirementAdequacy(regions, requirements)`)
1026
+ answers only "does the reference definition itself contain enough evidence
1027
+ to understand every selected requirement?" - never "does a runtime
1028
+ target/candidate exist" (that is Prompt 4/5's responsibility). It is its own
1029
+ small, reference-owned vocabulary (`REFERENCE_REQUIREMENT_ADEQUACY_STATES` =
1030
+ `'adequate' | 'partial' | 'inadequate'`, and exactly two reason codes,
1031
+ `no-selected-requirements` and `missing-reference-relationship-evidence`) -
1032
+ deliberately **not** a reuse of
1033
+ `boundedAgentContext.ts`'s `Adequacy`/`ADEQUACY_REASON_CODES`, which
1034
+ describe runtime-target/static-correlation concerns that do not exist at
1035
+ this stage; mislabeling reference adequacy as bounded-agent-context adequacy
1036
+ would conflate two genuinely different evidence domains. Zero selected
1037
+ requirements is explicitly `inadequate` (a region-rich, fully-valid
1038
+ reference is still not usable for a correction task until the user has
1039
+ actually selected what matters) - this is a documented product decision,
1040
+ not an oversight. The result is never a numeric score, always structured
1041
+ and inspectable, with reasons ordered deterministically by authored
1042
+ requirement position.
1043
+
1044
+ ```ts
1045
+ interface ReferenceRequirementAdequacy {
1046
+ status: 'adequate' | 'partial' | 'inadequate';
1047
+ totalRequirements: number;
1048
+ evaluableRequirements: number;
1049
+ unavailableRequirements: number;
1050
+ reasons: { code: 'no-selected-requirements' | 'missing-reference-relationship-evidence'; requirementId?: string; detail?: string }[];
1051
+ }
1052
+ ```
1053
+
1054
+ ## v0.7 Prompt 4 reference applicability and candidate-state compatibility
1055
+
1056
+ Released as `0.7.0`. Additive extension of the Prompt 1/2/3 contracts
1057
+ above: one new, optional `applicability?: ExternalReferenceApplicability`
1058
+ field on `ExternalReferenceArtifact` (both lifecycle variants), one new,
1059
+ optional `explicitState?: ExplicitStateDimensions` field on
1060
+ `ObservationArtifact.requestConfig`, and one new pure module,
1061
+ `domain/externalReferenceCompatibility.ts`, that answers a single question:
1062
+ "does this external reference describe the same frontend state as this
1063
+ candidate `ObservationArtifact`?" No schema version bump on either artifact
1064
+ - same reasoning as Prompt 2/3's additive fields.
1065
+
1066
+ **Central distinction**: this is page/state-level compatibility only -
1067
+ never geometry, never fidelity, never a visual/pixel comparison, and never
1068
+ region-to-runtime-target binding (Prompt 5). It answers "should a
1069
+ reference-vs-candidate geometry comparison even be attempted", not "does the
1070
+ candidate match the reference". Reference-evidence adequacy (Prompt 3) and
1071
+ reference/candidate compatibility (Prompt 4) are deliberately independent:
1072
+ a reference can be `adequate` (enough selected requirements to evaluate)
1073
+ while simultaneously `incomparable` against a given candidate (wrong
1074
+ viewport/theme/state), and vice versa - neither result constrains the
1075
+ other.
1076
+
1077
+ **State identity is always explicit, never inferred.** `theme`,
1078
+ `applicationState`, and `authenticatedState` are caller/configuration-
1079
+ supplied labels only. The observer never reads screenshot pixels, CSS, DOM
1080
+ classes/text, URLs, source code, filenames, accessibility labels,
1081
+ localStorage, or cookies to determine state - there is no automatic state
1082
+ detection anywhere in this codebase, and Prompt 4 does not add any. Labels
1083
+ are bounded opaque identities (`^[A-Za-z0-9_-]{1,64}$`, the same pattern
1084
+ already used for target names and region ids) compared by exact,
1085
+ case-sensitive string equality only - `"dark"` and `"one-dark"` are
1086
+ unrelated labels, never fuzzy-matched or normalized.
1087
+
1088
+ ```ts
1089
+ // domain/explicitState.ts - shared by both ObservationArtifact and ExternalReferenceArtifact
1090
+ type AuthenticatedState = 'authenticated' | 'unauthenticated'; // closed vocabulary - never a place for credentials/tokens/cookies/session ids
1091
+ interface ExplicitStateDimensions {
1092
+ theme?: string;
1093
+ applicationState?: string;
1094
+ authenticatedState?: AuthenticatedState;
1095
+ }
1096
+ // isValidExplicitStateDimensions requires at least one dimension declared and rejects any unsupported field -
1097
+ // this is a bounded, closed shape, never an arbitrary Record<string, unknown> metadata bag.
1098
+
1099
+ // domain/externalReferenceApplicability.ts
1100
+ interface ApplicableViewport { width: number; height: number } // CSS pixels, bounds [200, 3840] mirroring request.ts's own viewport bounds
1101
+ interface ExternalReferenceApplicability extends ExplicitStateDimensions {
1102
+ viewport?: ApplicableViewport;
1103
+ }
1104
+ ```
1105
+
1106
+ **Reference image size is never the same concept as applicable viewport.**
1107
+ `ExternalReferenceImageReference.width/height` (Prompt 1) describes the
1108
+ reference image's own pixel dimensions - a property of the image file,
1109
+ detected from its header bytes. `applicability.viewport` describes the
1110
+ CSS-pixel runtime viewport the design *represents* - a reference image may
1111
+ be captured at any resolution or device-pixel-ratio (e.g. a 1920x1080
1112
+ screenshot representing a 960x540 CSS-pixel layout at 2x DPR). Nothing in
1113
+ `externalReferenceApplicability.ts` reads or derives a viewport from image
1114
+ dimensions; `isValidExternalReferenceApplicability` is its own validator
1115
+ (not a reuse of `isValidExplicitStateDimensions`, whose "at least one
1116
+ dimension" rule would incorrectly reject a viewport-only applicability
1117
+ object).
1118
+
1119
+ **v0.4 comparability is reused, not duplicated.** `domain/comparison.ts`
1120
+ gained four additive reason codes (`viewport-unassessed`, `theme-mismatch`,
1121
+ `authenticated-state-mismatch`, `application-state-mismatch` - the
1122
+ `*-unassessed` codes for theme/authenticated-state/application-state
1123
+ already existed from v0.4) and two optional fields on `ComparabilityReason`
1124
+ (`referenceValue?: string`, `candidateValue?: string`, populated only for a
1125
+ mismatch reason). `domain/comparisonEngine.ts` gained one new exported pure
1126
+ helper, `assessOptionalComparabilityDimension(mismatchCode, unassessedCode,
1127
+ beforeValue, afterValue, mismatchMessage, unassessedMessage)`, extracted
1128
+ from - and now used by - both v0.4's own `evaluateComparability`
1129
+ (Observation-vs-Observation) and the new
1130
+ `evaluateReferenceCandidateCompatibility` (Reference-vs-Observation). The
1131
+ rule is identical either way: both values defined and equal -> no reason;
1132
+ both defined and different -> a `blocking` mismatch reason (with
1133
+ `referenceValue`/`candidateValue` populated); either value undefined ->
1134
+ an `unassessed` reason. This is a genuine, additive improvement to v0.4's
1135
+ own behavior: `evaluateComparability` now assesses theme/authenticated-
1136
+ state/application-state as matching or blocking-mismatched whenever *both*
1137
+ observations declare `requestConfig.explicitState`, rather than always
1138
+ reporting them unassessed - but every historical observation pair (and any
1139
+ pair where either side omits `explicitState`) retains the exact old
1140
+ unassessed-only behavior, verified by the frozen `evaluateComparability`
1141
+ regression test that predates this batch.
1142
+
1143
+ ```ts
1144
+ // domain/externalReferenceCompatibility.ts
1145
+ interface ReferenceCandidateCompatibilityResult {
1146
+ referenceId: string;
1147
+ referenceRequestId: string;
1148
+ candidateObservationId: string;
1149
+ candidateRequestId: string;
1150
+ compatibility: ComparabilityResult; // v0.4's own reused result type - state/reasons, never a boolean or a visual score
1151
+ }
1152
+ function evaluateReferenceCandidateCompatibility(reference: ExternalReferenceArtifact, candidate: ObservationArtifact): ReferenceCandidateCompatibilityResult;
1153
+ ```
1154
+
1155
+ Key rules:
1156
+
1157
+ - Pure and synchronous - no browser, no filesystem, no network, no target
1158
+ binding. Only `reference.applicability` and
1159
+ `candidate.requestConfig.viewport`/`candidate.requestConfig.explicitState`
1160
+ are consulted; reference regions/requirements are never read here (a
1161
+ distinct, separate concern - see Prompt 3 above).
1162
+ - A dimension the reference constrains but the candidate entirely omits
1163
+ (or vice versa) is `unassessed`, never treated as compatible-by-default
1164
+ and never fabricated as a mismatch - fail-closed, honest non-assessment.
1165
+ - A reference that declares no `applicability` at all produces a fully
1166
+ `unassessed` (never automatically `incomparable`, never automatically
1167
+ `comparable` beyond "no blocking reasons found") result across all four
1168
+ dimensions - Prompt 1/2/3 references remain fully usable, just
1169
+ unassessed for compatibility until applicability is authored.
1170
+ - No automatic persisted compatibility artifact. This is a pure
1171
+ programmatic result, produced on demand by an application/CLI caller
1172
+ that already holds both a reference and a candidate artifact - inventing
1173
+ a new persisted artifact kind for a value this cheap to recompute would
1174
+ add drift risk (a candidate/reference re-imported later could silently
1175
+ disagree with a stale persisted compatibility record) with no
1176
+ corresponding benefit; this may be revisited only if a later prompt's
1177
+ architecture proves persistence necessary.
1178
+ - Identity impact: `buildExternalReferenceRequestIdentity` gained a final
1179
+ optional `applicability` parameter (omitted, never `null`, when absent -
1180
+ byte-identical to Prompt 1/2/3 hashes for every call that doesn't supply
1181
+ it); `buildRequestIdentity` gained a final optional `explicitState`
1182
+ parameter with the identical omission convention. Neither identity
1183
+ function ever takes a file path.
1184
+ - CLI: `import-reference` gained an optional `--applicability-file
1185
+ <json-file>` (the raw, unwrapped applicability object - not a
1186
+ `{ "requirements": [...] }`-style wrapper, since applicability is a
1187
+ single object rather than a named list); `observe` gained an optional
1188
+ `--state-file <json-file>` (the raw, unwrapped `ExplicitStateDimensions`
1189
+ object). Both follow the existing `--scroll-scenario-file` convention
1190
+ exactly: relative paths resolve from the current working directory, the
1191
+ path itself is never persisted or included in any identity, and CLI code
1192
+ owns only flag syntax/file reading/JSON parsing/object-root validation -
1193
+ all semantic validation happens in the domain layer.
1194
+
1195
+ ## v0.7 Prompt 5 explicit reference-region <-> runtime-target binding
1196
+
1197
+ Released as `0.7.0`. One new pure domain module,
1198
+ `domain/externalReferenceRuntimeBinding.ts`, answering "which stable
1199
+ observer runtime target, if any, does this candidate observation resolve
1200
+ for each explicitly declared reference region?" No new field is added to
1201
+ either `ExternalReferenceArtifact` or `ObservationArtifact` - both remain
1202
+ exactly as Prompt 4 left them - and no schema version bump on either.
1203
+
1204
+ **Two identity domains, kept strictly separate.** A binding declaration
1205
+ names a Prompt 2 `ReferenceRegion.id` and a v0.2 `NamedTarget.name` (the
1206
+ stable observer runtime target identity established since v0.2 - never a
1207
+ CSS selector, DOM node handle, source file, React component name, or
1208
+ my-dev-kit node id). These two strings living in the same textual namespace
1209
+ never implies a binding - a region id `"header"` and a target name
1210
+ `"header"` bind to each other only because of an explicit declaration, not
1211
+ because the strings match (verified by a dedicated test: the same
1212
+ observation with and without the explicit declaration produces `bound`
1213
+ only in the former case).
1214
+
1215
+ ```ts
1216
+ // domain/externalReferenceRuntimeBinding.ts
1217
+ interface ReferenceRuntimeBindingDeclaration {
1218
+ referenceRegion: string; // Prompt 2 ReferenceRegion.id
1219
+ runtimeTarget: string; // v0.2 NamedTarget.name
1220
+ }
1221
+
1222
+ const REFERENCE_RUNTIME_BINDING_STATUSES = ['bound', 'ambiguous', 'unavailable'] as const;
1223
+
1224
+ interface ReferenceRuntimeBindingResult {
1225
+ referenceRegion: string;
1226
+ runtimeTarget: string;
1227
+ status: 'bound' | 'ambiguous' | 'unavailable';
1228
+ reasonCode?: 'runtime-target-not-configured' | 'runtime-target-not-found' | 'runtime-target-ambiguous' | 'runtime-target-evidence-unavailable';
1229
+ detail: string;
1230
+ targetResolutionStatus?: TargetSelectionStatus; // v0.2's own resolution status, when evidence for it exists
1231
+ targetVisible?: boolean; // provenance only - never affects status
1232
+ }
1233
+
1234
+ interface ReferenceRuntimeBindingEvaluation {
1235
+ referenceId: string;
1236
+ referenceRequestId: string;
1237
+ candidateObservationId: string;
1238
+ candidateRequestId: string;
1239
+ compatibility: ComparabilityResult; // reused verbatim from v0.7 Prompt 4
1240
+ bindings: ReferenceRuntimeBindingResult[]; // empty exactly when compatibility.state === 'incomparable'
1241
+ }
1242
+
1243
+ function evaluateReferenceRuntimeBindings(
1244
+ reference: ExternalReferenceArtifact,
1245
+ candidate: ObservationArtifact,
1246
+ declarations: readonly ReferenceRuntimeBindingDeclaration[],
1247
+ ): { ok: true; evaluation: ReferenceRuntimeBindingEvaluation } | { ok: false; reason: string };
1248
+ ```
1249
+
1250
+ Key rules:
1251
+
1252
+ - **Explicit, never inferred.** A binding declaration is user/configuration
1253
+ input asserting a conceptual correspondence; this module never discovers
1254
+ it from screenshot geometry, matching names, matching text, or source
1255
+ code. There is no automatic-matching algorithm anywhere in this module.
1256
+ - **Reuses, never duplicates.** The compatibility gate reuses
1257
+ `evaluateReferenceCandidateCompatibility` (Prompt 4) verbatim - viewport/
1258
+ theme/application-state/authenticated-state comparison logic is never
1259
+ re-implemented here. Runtime-target resolution reuses `targetPresence`
1260
+ (v0.4 `comparisonEngine.ts`, now additively exported alongside
1261
+ `assessOptionalComparabilityDimension`) - the exact same "how do I read a
1262
+ `TargetEvidenceRecord`'s resolution" rule v0.4's own before/after target
1263
+ comparison already uses. No second target resolver, no browser launch, no
1264
+ Chromium query, no selector evaluation, no live-DOM inspection - this
1265
+ module consumes only an already-captured `ObservationArtifact`'s
1266
+ `requestConfig.targets`/`targetEvidence`.
1267
+ - **Compatibility gates before evaluation, structurally.** If
1268
+ `evaluateReferenceCandidateCompatibility` reports `incomparable`,
1269
+ `bindings` is the empty array and the caller reads the reason from the
1270
+ embedded `compatibility` field - there is no binding-local "incompatible"
1271
+ status; Prompt 4's compatibility result is represented exactly once, not
1272
+ duplicated into a parallel vocabulary.
1273
+ - **Two-layer validation.** Reference-region existence, declaration shape,
1274
+ bounds, and duplicate/conflict rules are validated structurally against
1275
+ the `ExternalReferenceArtifact` alone (`isValidReferenceRuntimeBindingDeclarations`)
1276
+ - independent of any candidate, mirroring Prompt 3's "unknown region
1277
+ reference is a structural validation failure" precedent exactly: a
1278
+ declaration naming a nonexistent reference region, or any reference with
1279
+ no `regions` declared at all, fails the whole evaluation closed before a
1280
+ candidate is even considered. Runtime-target availability, by contrast,
1281
+ is evaluated per-candidate inside `evaluateReferenceRuntimeBindings`
1282
+ itself, since the same declaration can be `bound` against one candidate
1283
+ and `unavailable` against another.
1284
+ - **Duplicate/conflicting-declaration rule** (mirrors Prompt 3's
1285
+ requirement-subject uniqueness rule): no two declarations may name the
1286
+ same `referenceRegion` (case-insensitively), whether they agree on
1287
+ `runtimeTarget` (an exact duplicate) or disagree (a conflict) - both fail
1288
+ the same way, never silently resolved by keeping the first. The reverse -
1289
+ several distinct reference regions naming the same `runtimeTarget` - is
1290
+ deliberately allowed (e.g. two design sub-regions legitimately
1291
+ corresponding to one runtime container element).
1292
+ - **Target-resolution-state handling.** `targetPresence`'s four outcomes
1293
+ map onto binding status as: `matched` -> `bound`; `ambiguous` -> `ambiguous`
1294
+ (the candidate's own configured target resolved ambiguously - never
1295
+ reported bound even though its stable name exists); `not-found` ->
1296
+ `unavailable` (`runtime-target-not-found` - the target was configured but
1297
+ the resolver found nothing on the page); no usable resolution evidence at
1298
+ all -> `unavailable` (`runtime-target-evidence-unavailable`). A declared
1299
+ `runtimeTarget` that was never part of the candidate's configured target
1300
+ set at all is a fifth, CLI/config-boundary-only outcome -> `unavailable`
1301
+ (`runtime-target-not-configured`) - never a dynamic page search.
1302
+ - **Hidden-target decision.** A uniquely resolved (`matched`) but hidden
1303
+ target is still reported `bound` - visibility never changes `status`.
1304
+ `targetVisible` (from the existing `TargetVisibility` evidence, when
1305
+ available) is carried as provenance only. Binding identity (does a stable
1306
+ correspondence exist) and later fidelity evaluability (can this evidence
1307
+ actually be used to check the design) are treated as distinct questions;
1308
+ this prompt answers only the former.
1309
+ - **Not every region needs a binding.** `isValidReferenceRuntimeBindingDeclarations`
1310
+ never requires full region coverage - a reference may have regions no
1311
+ declaration names at all (they simply have no bound runtime target for
1312
+ this candidate). This is not a completeness gate; Prompt 6 (or later) may
1313
+ add one for the regions that selected requirements actually need.
1314
+ - **No persisted artifact family.** `evaluateReferenceRuntimeBindings` is a
1315
+ pure, on-demand function over an already-persisted reference, an
1316
+ already-persisted candidate observation, and an in-memory declaration
1317
+ collection. No `ExternalReferenceBindingArtifact` (or equivalent) is
1318
+ introduced - the same "cheap to recompute, persisting invites drift"
1319
+ reasoning Prompt 4 already applied to its own compatibility result.
1320
+ Neither the reference nor the observation artifact is ever rewritten to
1321
+ carry a binding result: a design reference may later be evaluated against
1322
+ several different candidates, and one observation may be evaluated
1323
+ against several different references, so binding is kept as downstream,
1324
+ candidate-specific, reference-specific derived evidence rather than
1325
+ mutating either immutable source artifact.
1326
+ - **No new identity function.** Unlike `buildRequestIdentity`/
1327
+ `buildExternalReferenceRequestIdentity`, no hash-based logical identity is
1328
+ computed for a binding declaration or its evaluated result - there is no
1329
+ persistence and no cross-document reference-by-id need yet (mirroring
1330
+ Prompt 4's `ReferenceCandidateCompatibilityResult`, which took the same
1331
+ approach). Provenance is instead carried directly as plain fields
1332
+ (`referenceId`, `referenceRequestId`, `candidateObservationId`,
1333
+ `candidateRequestId`, plus each result's own `referenceRegion`/
1334
+ `runtimeTarget`) - already deterministic, already sufficient for a caller
1335
+ to trace every result back to its inputs, without inventing a fifth
1336
+ identity-hashing convention for a value this prompt does not persist.
1337
+ - **Deterministic ordering.** `bindings` preserves authored declaration
1338
+ order (mirroring the "authored order is semantic" convention already used
1339
+ for regions/requirements) rather than sorting by any derived key.
1340
+ - **Bounded.** `MAX_REFERENCE_RUNTIME_BINDINGS` (20) caps the declaration
1341
+ collection, mirroring `MAX_REFERENCE_REGIONS`.
1342
+ - **No public CLI surface yet.** Only the programmatic
1343
+ `evaluateReferenceRuntimeBindings`/`isValidReferenceRuntimeBindingDeclarations`
1344
+ functions are exported. A standalone CLI command was deliberately not
1345
+ added merely for symmetry with `import-reference`/`observe`; Prompt 6
1346
+ (structured fidelity evaluation) is expected to become the first concrete
1347
+ consumer and public-surface owner for this capability.
1348
+
1349
+ ## v0.7 Prompt 6 structured reference-vs-candidate fidelity evaluation
1350
+
1351
+ Released as `0.7.0`. One new pure domain module,
1352
+ `domain/externalReferenceFidelity.ts`, and its CLI-facing counterpart,
1353
+ `application/referenceFidelityEvaluationService.ts` plus the new
1354
+ `evaluate-reference-fidelity` CLI command - the first point in this whole
1355
+ v0.7 stack where a reference's authored expectation is actually compared
1356
+ against live candidate evidence. No new artifact field, no schema version
1357
+ bump: this prompt reuses Prompt 1-5's artifacts and result types entirely.
1358
+
1359
+ ```ts
1360
+ // domain/externalReferenceFidelity.ts
1361
+ const REFERENCE_REQUIREMENT_FIDELITY_STATUSES = ['pass', 'fail', 'unavailable'] as const;
1362
+ const REFERENCE_FIDELITY_STATES = ['not-evaluated', 'pass', 'fail'] as const;
1363
+ const REFERENCE_FIDELITY_BLOCK_REASONS = ['reference-inadequate', 'incompatible'] as const;
1364
+
1365
+ interface ReferenceRequirementFidelityResult {
1366
+ requirementId: string;
1367
+ category: AuthoredChangeScopeCategory;
1368
+ expectedDependentMode?: ExpectedDependentMode;
1369
+ subject: ReferenceRequirementSubject;
1370
+ boundRuntimeTargets: string[];
1371
+ status: 'pass' | 'fail' | 'unavailable';
1372
+ reasonCode?: 'reference-evidence-unavailable' | 'reference-relationship-not-exhibited' | 'binding-unavailable' | 'candidate-evidence-unavailable' | 'coordinate-mapping-unavailable'; // present iff status === 'unavailable'
1373
+ detail?: string;
1374
+ // region-property/region-measurement subjects only:
1375
+ referenceValue?: number; // reference-image pixels
1376
+ candidateRawValue?: number; // CSS pixels, as captured
1377
+ candidateValue?: number; // candidateRawValue converted into reference-image-pixel space
1378
+ delta?: number; // candidateValue - referenceValue
1379
+ tolerance?: ReferenceRequirementTolerance;
1380
+ // region-relationship subjects only:
1381
+ expectedRelationship?: PairwiseRelationshipKind;
1382
+ actualRelationship?: PairwiseRelationshipKind;
1383
+ }
1384
+
1385
+ interface ReferenceCandidateFidelityEvaluation {
1386
+ referenceId: string;
1387
+ referenceRequestId: string;
1388
+ candidateObservationId: string;
1389
+ candidateRequestId: string;
1390
+ adequacy: ReferenceRequirementAdequacy; // reused verbatim from Prompt 3
1391
+ compatibility?: ComparabilityResult; // reused verbatim from Prompt 4; absent only when adequacy itself is inadequate
1392
+ bindings?: ReferenceRuntimeBindingEvaluation; // reused verbatim from Prompt 5; absent when an earlier gate blocked
1393
+ state: 'not-evaluated' | 'pass' | 'fail';
1394
+ blockedBy?: 'reference-inadequate' | 'incompatible'; // present iff state === 'not-evaluated'
1395
+ requirementResults: ReferenceRequirementFidelityResult[]; // empty iff state === 'not-evaluated'
1396
+ }
1397
+
1398
+ function evaluateReferenceCandidateFidelity(
1399
+ reference: ExternalReferenceArtifact,
1400
+ candidate: ObservationArtifact,
1401
+ bindingDeclarations: readonly ReferenceRuntimeBindingDeclaration[],
1402
+ options?: { geometryTolerancePx?: number },
1403
+ ): { ok: true; evaluation: ReferenceCandidateFidelityEvaluation } | { ok: false; reason: string };
1404
+ ```
1405
+
1406
+ **Result vocabulary.** `pass`/`fail`/`unavailable` is reused from v0.5's
1407
+ `CLAUSE_RESULT_STATUSES` shape (the same honest three-state idea: a result
1408
+ either satisfies its condition, fails it, or cannot be evaluated - never a
1409
+ score) but is its own independently-owned constant, deliberately excluding
1410
+ v0.5's fourth member, `'conflict'` - Prompt 6 has no cross-requirement
1411
+ authoring-conflict concept (each requirement is evaluated independently
1412
+ against its own subject), so reusing `conflict` would invite a status this
1413
+ prompt can never actually produce.
1414
+
1415
+ **Evaluation order (frozen, never reordered):** reference structural
1416
+ validation -> candidate structural validation -> binding-declaration
1417
+ structural validation -> Prompt 3 reference adequacy -> Prompt 4
1418
+ compatibility -> Prompt 5 binding evaluation -> per-requirement candidate-
1419
+ evidence/coordinate-mapping checks -> per-requirement tolerance/
1420
+ relationship comparison -> overall result. The first three (structural
1421
+ validation) failures return `{ ok: false, reason }` - a caller/config error,
1422
+ never a fidelity outcome. The next two (adequacy `inadequate`, compatibility
1423
+ `incomparable`) short-circuit to `state: 'not-evaluated'` with an empty
1424
+ `requirementResults` - an earlier blocking gate never lets an ordinary
1425
+ PASS/FAIL requirement set get fabricated past it. `adequacy` "partial" (some,
1426
+ but not all, authored requirements individually unavailable) does **not**
1427
+ block evaluation - it proceeds normally, and the individual unavailable
1428
+ reference-side requirements simply also report `unavailable` at the
1429
+ per-requirement level (their own reference-evidence problem, re-derived
1430
+ identically by `evaluateOneRequirement`, not looked up from the adequacy
1431
+ result).
1432
+
1433
+ **Coordinate mapping - the central problem this prompt solves.** Prompt 2
1434
+ regions and Prompt 3 tolerances are authored in reference-image pixels;
1435
+ `ObservationArtifact` target geometry is CSS pixels. This module establishes
1436
+ exactly one explicit, deterministic scale from `reference.applicability.viewport`
1437
+ (the CSS-pixel runtime viewport, Prompt 4) and the reference image's own
1438
+ pixel dimensions (Prompt 1) - `scaleX = imageWidth / viewportWidth`,
1439
+ `scaleY = imageHeight / viewportHeight` - and converts every candidate
1440
+ measurement into reference-image-pixel space before comparing it against a
1441
+ Prompt 3 tolerance. It never assumes 1 reference-image pixel equals 1 CSS
1442
+ pixel, and it never performs cropping, offset, rotation, or perspective
1443
+ registration - only a deliberately bounded full-frame mapping. `scaleX`/
1444
+ `scaleY` must agree within a small, independently-owned coordinate-mapping-
1445
+ validity tolerance (1% relative, never a user-authored design tolerance) or
1446
+ the mapping is rejected outright; a reference with no applicable viewport at
1447
+ all likewise has no mapping. Either way, every numeric (`region-property`/
1448
+ `region-measurement`) requirement becomes `unavailable`/`coordinate-mapping-unavailable`
1449
+ - categorical `region-relationship` requirements are unaffected (they never
1450
+ need a scale). Horizontal fields (`x`/`width`/`right`/`centerX` and the
1451
+ horizontal measurements) always scale by `scaleX`; vertical fields (`y`/
1452
+ `height`/`bottom`/`centerY` and the vertical measurements) always scale by
1453
+ `scaleY` - this falls out automatically from converting a full
1454
+ `TargetGeometry` into a `ReferenceRegionGeometry`-shaped value per axis,
1455
+ never a hand-picked per-property axis table.
1456
+
1457
+ **Tolerance is reused exactly, never redefined.** A single rule -
1458
+ `abs(delta) <= allowedAmount` - covers all three Prompt 3 tolerance kinds:
1459
+ `exact` is simply the zero-tolerance case (`allowedAmount = 0`);
1460
+ `absolute-reference-px` uses its authored `amount` directly (already in
1461
+ reference-image pixels); `percent`'s denominator is `Math.abs(referenceValue)`,
1462
+ mirroring v0.5's own `toleranceToPx` "may vary by up to N%" convention
1463
+ exactly (independently reimplemented in reference-image-pixel units, never
1464
+ imported - `frontendContractEvaluation.ts`'s `ContractTolerance` is a
1465
+ different, CSS-pixel-implicit unit). No hidden epsilon is added anywhere;
1466
+ subpixel precision is preserved through to the final comparison, so a
1467
+ tolerance-boundary value (e.g. delta exactly equal to the allowed amount)
1468
+ passes and one unit past it fails, exactly as authored.
1469
+
1470
+ **Region-property evaluation** reads `TargetGeometry` from the bound
1471
+ target's `targetEvidence` entry, converts it into a `ReferenceRegionGeometry`-
1472
+ shaped value (adding `centerX`/`centerY`, computed identically to
1473
+ `deriveReferenceRegionGeometry`) both raw (CSS) and scaled (reference-image
1474
+ pixels), and reads `[subject.property]` off each - `candidateRawValue`
1475
+ (CSS) and `candidateValue` (reference-image pixels) are both reported.
1476
+
1477
+ **Region-measurement evaluation** converts *both* bound targets' geometries
1478
+ the same way and calls the existing `deriveReferenceRequirementMeasurement`
1479
+ (Prompt 3) on the converted geometries directly - reusing Prompt 3's exact
1480
+ gap/delta formulas rather than reimplementing a parallel "runtime version"
1481
+ of them, and never inventing a generic geometry expression language. A
1482
+ geometrically-undefined gap (the two targets overlap on the relevant axis)
1483
+ is `unavailable`, mirroring Prompt 3's own reference-side treatment of the
1484
+ identical situation.
1485
+
1486
+ **Region-relationship evaluation** first confirms the reference itself
1487
+ actually exhibits its own selected relationship (`deriveReferenceRequirementExpectation`'s
1488
+ `matches` field) - if not, the result is `unavailable`/
1489
+ `reference-relationship-not-exhibited` (a reference-authoring problem, never
1490
+ a candidate `fail`). It then resolves both bound targets and calls the
1491
+ canonical `deriveLayoutRelationships` (v0.4) over the *whole* candidate
1492
+ observation - never a second, parallel relationship formula - and looks up
1493
+ the pairwise record for the bound target pair **scoped to the exact
1494
+ requested relationship family** (an independently-owned, third duplicate of
1495
+ the same `RELATIONSHIP_FAMILY_GROUPS` shape already used by
1496
+ `frontendContractEvaluation.ts` and `externalReferenceRequirements.ts` -
1497
+ this is the same real bug class Prompt 3 fixed: matching the first record
1498
+ for a target pair regardless of family would silently compare against the
1499
+ wrong relationship kind). A record only derivable in the reversed target
1500
+ order is `unavailable`, never auto-flipped - identical to Prompt 3's own
1501
+ reference-side handling of the same situation. `pass` requires the
1502
+ candidate's actual relationship kind to equal the requirement's authored
1503
+ `relationship` exactly.
1504
+
1505
+ **Binding gate.** Every subject's dependent reference region(s) must have a
1506
+ `bound` (never `ambiguous`/`unavailable`, and never simply absent from the
1507
+ supplied declarations) Prompt 5 binding result, or the requirement is
1508
+ `unavailable`/`binding-unavailable` - this module never guesses another
1509
+ target and never auto-binds based on geometry or names. A `bound` target
1510
+ that is not visible (`TargetVisibility.visible !== true`, including when
1511
+ visibility evidence itself is unavailable) is treated as having no usable
1512
+ geometry - `unavailable`/`candidate-evidence-unavailable` - preserving the
1513
+ distinction between "binding succeeded" (Prompt 5's question) and "this
1514
+ evidence is usable for fidelity evaluation" (this prompt's question): a
1515
+ hidden-but-uniquely-resolved target still has a stable identity, but its
1516
+ geometry is never treated as meaningful for a numeric/relationship
1517
+ comparison.
1518
+
1519
+ **Categories are preserved, never given different PASS/FAIL rules.**
1520
+ `category`/`expectedDependentMode` are carried through to each result as
1521
+ provenance only; Prompt 3 never implemented a `required`-vs-`permitted`
1522
+ directional evaluation difference for its own expectation/adequacy
1523
+ derivation (unlike v0.5's runtime-directional contract clauses), so Prompt 6
1524
+ does not invent one now - every requirement in the reference's authored
1525
+ collection is evaluated by the identical rule and counts identically toward
1526
+ the overall result, regardless of category.
1527
+
1528
+ **Overall fidelity result.** For an evaluated (non-blocked) pair, `state`
1529
+ is `'pass'` only when every requirement result is `'pass'`; any `'fail'` or
1530
+ `'unavailable'` result forces `state: 'fail'` - there is no meaningful third
1531
+ overall bucket once evaluation has actually run, since "some/all
1532
+ unavailable" and "some/all fail" both equally mean "not every selected
1533
+ requirement is confirmed satisfied." A reference with zero selected
1534
+ requirements never reaches this stage at all - it is `inadequate` (Prompt
1535
+ 3's own zero-requirements rule) and therefore `not-evaluated`, never a
1536
+ meaningless `pass`.
1537
+
1538
+ **Persistence decision: none.** `evaluateReferenceCandidateFidelity` (and
1539
+ its CLI-facing wrapper, `evaluateReferenceCandidateFidelityFromArtifactRoots`)
1540
+ is a pure, on-demand function over already-persisted/in-memory evidence -
1541
+ no new `ExternalReferenceFidelityEvaluationArtifact` (or equivalent) is
1542
+ introduced. Rationale, identical to Prompt 4/5's own precedent: the result
1543
+ is cheap to recompute deterministically from its inputs (a reference, a
1544
+ candidate, and a caller-supplied binding-declaration collection), and
1545
+ persisting it would invite drift with no corresponding benefit at this
1546
+ stage; this may be revisited only if Prompt 7's architecture proves
1547
+ persistence necessary.
1548
+
1549
+ **CLI**: `evaluate-reference-fidelity --reference <root> --candidate <root>
1550
+ [--bindings-file <json-file>] [--enforce]` - the CLI surface Prompt 5
1551
+ deliberately deferred. `--bindings-file` follows the exact
1552
+ `--requirements-file`/`--regions-file` wrapped-object convention
1553
+ (`{ "bindings": [...] }`); CLI code owns only flag syntax/file reading/JSON
1554
+ parsing/root-shape validation, with every binding-declaration rule staying
1555
+ owned by `isValidReferenceRuntimeBindingDeclarations`. `--enforce` mirrors
1556
+ `evaluate-contract`'s exact precedent: it changes only the process exit
1557
+ status for an already-computed `state: 'fail'` result, never the printed
1558
+ content - and has no effect on `not-evaluated`, which always exits 0 (a
1559
+ compatibility/adequacy blocker is a successful, structured, honest
1560
+ non-evaluation, never an execution error and never a design mismatch).
1561
+ Persists nothing; there is no `--output` flag.
1562
+
1563
+ ## v0.7 Prompt 7 bounded reference-fidelity projection and v0.6 bounded-agent-context integration
1564
+
1565
+ Released as `0.7.0`. Additive extension of the v0.6 bounded-agent-context
1566
+ contract above and of the v0.7 Prompt 6 fidelity evaluator - no new bounded-
1567
+ context artifact family, no second visual-context system, no schema version
1568
+ bump (`BOUNDED_AGENT_CONTEXT_SCHEMA_VERSION` stays `1.0.0`, following the
1569
+ exact precedent already set when `correlations?` was added in v0.6 Batch 3).
1570
+
1571
+ **Chosen integration owner.** `projectBoundedAgentContext` itself gains one
1572
+ new optional input (`fidelity?: ReferenceCandidateFidelityEvaluation`, plus
1573
+ `fidelityRequired?: boolean`) rather than a separate `VisualAgentContext`/
1574
+ `VisualPromptPacket`/`ReferencePromptBuilder`. This was chosen over a pure
1575
+ post-hoc "attach" step (the shape `attachRuntimeStaticCorrelations` uses)
1576
+ because fidelity-relevant runtime targets must compete fairly for
1577
+ `MAX_RUNTIME_TARGETS` capacity and receive the exact same geometry/
1578
+ visibility/screenshot assembly contract-clause-derived targets already get -
1579
+ an attach-only step run after target allocation could never produce that. A
1580
+ new pure module, `domain/referenceFidelityProjection.ts`
1581
+ (`projectReferenceFidelity`), derives the bounded, prioritized fidelity
1582
+ content plus the target-id/omission/truncation contributions
1583
+ `projectBoundedAgentContext` folds into its own existing pipeline - it is
1584
+ not a second fidelity-evaluation engine, only a selection over Prompt 6's
1585
+ already-computed result.
1586
+
1587
+ ```ts
1588
+ // domain/boundedAgentContext.ts - additive
1589
+ interface BoundedAgentContextSourceReferences {
1590
+ // ...unchanged fields...
1591
+ referenceId?: string; // new, optional
1592
+ referenceRequestId?: string; // new, optional
1593
+ }
1594
+
1595
+ const MAX_FIDELITY_MISMATCHES = 15;
1596
+ const MAX_FIDELITY_PROTECTED_CONTEXT = 10; // reuses MAX_RELATIONSHIP_EVIDENCE_PER_TARGET's value
1597
+
1598
+ interface BoundedReferenceFidelityProjection {
1599
+ referenceId: string;
1600
+ referenceRequestId: string;
1601
+ candidateObservationId: string;
1602
+ candidateRequestId: string;
1603
+ adequacy: ReferenceRequirementAdequacy; // reused verbatim from Prompt 3
1604
+ compatibility?: ComparabilityResult; // reused verbatim from Prompt 4
1605
+ state: ReferenceFidelityState; // reused verbatim from Prompt 6
1606
+ blockedBy?: ReferenceFidelityBlockReason; // reused verbatim from Prompt 6
1607
+ mismatches: ReferenceRequirementFidelityResult[]; // bounded, prioritized non-pass requirements (Prompt 6 type, unmodified)
1608
+ protectedContext: ReferenceRequirementFidelityResult[]; // bounded passing protected/preserved requirements, as "do not break this" context
1609
+ }
1610
+
1611
+ interface BoundedAgentContextArtifact {
1612
+ // ...unchanged fields...
1613
+ fidelity?: BoundedReferenceFidelityProjection; // new, optional - mirrors `correlations?`'s own additive precedent exactly
1614
+ }
1615
+ ```
1616
+
1617
+ **Selection policy** (`domain/referenceFidelityProjection.ts#projectReferenceFidelity`):
1618
+ only Prompt 6's non-`pass` requirement results are ever candidates for
1619
+ `mismatches` - passing requirements are never dumped by default, satisfying
1620
+ this prompt's "bounded coding-agent use" design goal. Each candidate is
1621
+ classified into a tier by its authored category/mode, reusing
1622
+ `boundedAgentContextProjection.ts#clauseTier`'s exact rule (duplicated, not
1623
+ imported, per this repository's established per-module small-helper
1624
+ convention - never a reference-specific protected/preserved taxonomy):
1625
+ `protected`/`preserved` are always `required`; `expected-dependent` is
1626
+ `required` only in `'required'` mode; `requested` and `expected-dependent`/
1627
+ `'permitted'` are `optional`.
1628
+
1629
+ **Priority policy**: 1) `fail` + `required` tier, 2) `unavailable` +
1630
+ `required` tier, 3) any other non-`pass` (optional-tier) result. Within one
1631
+ priority class, Prompt 6's own authored requirement order is preserved (a
1632
+ stable sort by priority rank only) - never re-ranked by an opaque score.
1633
+ The final `mismatches` array is reported in priority order (highest first),
1634
+ not restored to authored order, since the whole point of prioritization is
1635
+ that the most actionable evidence appears first when the set is large.
1636
+
1637
+ **Cap values**: `MAX_FIDELITY_MISMATCHES = 15` and
1638
+ `MAX_FIDELITY_PROTECTED_CONTEXT = 10` (reusing
1639
+ `MAX_RELATIONSHIP_EVIDENCE_PER_TARGET`'s value) - both judgment-call bounds
1640
+ in the same spirit as v0.6 Batch 1's own frozen caps (no measured fixture
1641
+ corpus exists yet for either concept).
1642
+
1643
+ **Omission/truncation behavior**: reuses `OmissionRecord`/`TruncationRecord`
1644
+ wholesale, no second reporting model. When mismatches exceed the cap, a
1645
+ `{subject: 'fidelity-mismatches', limit, actualCount, required}` truncation
1646
+ is recorded, plus one `{subject: 'fidelity-mismatch:<requirementId>',
1647
+ reason: 'required-evidence-lost-by-bound', required: true}` omission for
1648
+ *each* dropped required-tier mismatch (optional-tier drops are truncated
1649
+ but never separately omitted as "required loss", since they were never
1650
+ required). `protectedContext` truncation is always `required: false` - it
1651
+ is confirmatory/passing context, never a design-fidelity failure. These
1652
+ records are folded into `projectBoundedAgentContext`'s own `omissions`/
1653
+ `truncations` arrays *before* its existing aggregate `capOmissions`/
1654
+ `capTruncations` calls and its existing adequacy computation run - fidelity
1655
+ loss is never a separate adequacy code path, it simply participates in the
1656
+ exact same `anyRequiredLoss`/`anyOptionalLoss` rule every other evidence
1657
+ source already uses.
1658
+
1659
+ **Adequacy behavior**: a `not-evaluated` fidelity (blocked by Prompt 6's own
1660
+ `reference-inadequate`/`incompatible` gates) is never converted into "no
1661
+ problems" - `projectReferenceFidelity` records an explicit
1662
+ `{subject: 'fidelity', reason: 'unsupported-or-unavailable', required,
1663
+ detail}` omission, where `required` defaults to `true` (supplying a
1664
+ fidelity evaluation to be projected at all is itself the signal that the
1665
+ task depends on it, mirroring `CorrelationTargetInput.required`'s existing
1666
+ v0.6 convention - callers who want fidelity as purely incidental context set
1667
+ `fidelityRequired: false`). A `required: true` fidelity omission, folded
1668
+ into the existing adequacy computation, prevents `adequacy.state` from
1669
+ remaining `'adequate'` (it becomes `'partial'`, or `'inadequate'` when
1670
+ combined with other required loss reaching the existing threshold) - it is
1671
+ never silently ignored. A `required: false` omission can degrade adequacy
1672
+ to at most `'partial'`, per v0.6's own pre-existing "optional-only loss
1673
+ never means inadequate" rule - unchanged, not redefined. A `pass` fidelity
1674
+ result contributes no omissions/truncations at all and never degrades
1675
+ adequacy.
1676
+
1677
+ **Not-evaluated fidelity behavior**: preserved exactly as Prompt 6 reported
1678
+ it - `fidelity.state`/`fidelity.blockedBy` on the output artifact are a
1679
+ direct pass-through of Prompt 6's own values, with `mismatches`/
1680
+ `protectedContext` both empty (there is nothing to select from an empty
1681
+ `requirementResults`).
1682
+
1683
+ **Per-target organization**: every fidelity mismatch's `boundRuntimeTargets`
1684
+ (Prompt 5/6's own field, never truncated) becomes a required- or permitted-
1685
+ tier addition to `projectBoundedAgentContext`'s existing target-id sets,
1686
+ so those runtime targets receive full `BoundedRuntimeTargetProjection`
1687
+ treatment (geometry/visibility/overflow/scrollOwner/screenshotRef) through
1688
+ the exact existing assembly code - no duplicated target-projection logic.
1689
+ Reference regions are never used as a correlation or target-selection key;
1690
+ only the already-bound stable v0.2 runtime target ids are.
1691
+
1692
+ **Multi-target relationship representation**: a `region-relationship`
1693
+ mismatch's `boundRuntimeTargets` array (already carrying both bound
1694
+ targets, from Prompt 6) is used as-is - both targets are added to the
1695
+ required/permitted set, so both appear in `targets`. Nothing collapses a
1696
+ two-target relationship failure onto a single target.
1697
+
1698
+ **Static-correlation reuse**: entirely unchanged. `deriveRuntimeStaticCorrelations`/
1699
+ `attachRuntimeStaticCorrelations` are not modified, not called from within
1700
+ this prompt's new code, and remain the caller's own separate step -
1701
+ `BoundedRuntimeTargetProjection.targetId`/`RuntimeStaticCorrelationRecord.runtimeTargetId`
1702
+ already share the same stable v0.2 identity a fidelity mismatch's
1703
+ `boundRuntimeTargets` also uses, so a caller (Prompt 8) joins fidelity,
1704
+ target, and correlation evidence by that one shared id without this module
1705
+ ever needing to read source, run my-dev-kit, or choose among ambiguous
1706
+ candidates itself.
1707
+
1708
+ **Ambiguous/unavailable correlation behavior**: unaffected - a
1709
+ `RuntimeStaticCorrelationRecord` with `status: 'ambiguous'` continues to
1710
+ preserve every competing candidate (v0.6's own frozen invariant,
1711
+ untouched), and `status: 'unavailable'` never causes a fidelity mismatch
1712
+ for that same runtime target to be dropped - the two evidence kinds
1713
+ (runtime fidelity, static correlation) are attached independently and
1714
+ neither erases the other.
1715
+
1716
+ **Provenance**: every included mismatch remains traceable to the reference
1717
+ (`sources.referenceId`/`referenceRequestId`, new), the requirement
1718
+ (`requirementId`, `category`, `subject` - naming its reference region(s)),
1719
+ the Prompt 5 binding (`boundRuntimeTargets`), the candidate
1720
+ (`sources.observationIds`), and the full Prompt 6 evidence
1721
+ (`referenceValue`/`candidateRawValue`/`candidateValue`/`delta`/`tolerance`
1722
+ or `expectedRelationship`/`actualRelationship`) - nothing is replaced by a
1723
+ prose-only summary. No raw image bytes are ever embedded (fidelity carries
1724
+ only identifiers and numeric/categorical evidence, never pixels), and no
1725
+ source-ownership field (`sourceOwner`/`sourceFile`/`component`/`symbol`/
1726
+ `causedBy`) is ever produced - Prompt 7 stops at the runtime target exactly
1727
+ as Prompt 6 did; v0.6's own, unmodified static correlation is the only
1728
+ source-adjacent evidence this context ever carries, and it remains
1729
+ evidence, never edit authorization.
1730
+
1731
+ **Identity impact**: `buildBoundedAgentContextRequestIdentity` gained a
1732
+ final optional `fidelity?: unknown` parameter - omitted (never `null`) from
1733
+ the hashed semantic view when absent, so every pre-Prompt-7 call site keeps
1734
+ producing its exact byte-identical hash (verified by a frozen-vector-style
1735
+ regression test). When present, the caller's already-derived, bounded
1736
+ `BoundedReferenceFidelityProjection` (not the raw Prompt 6 evaluation) is
1737
+ hashed, so identity changes exactly when the content a caller would
1738
+ actually receive changes - never merely because an unselected, dropped
1739
+ requirement result changed somewhere upstream. `sources.referenceId`/
1740
+ `referenceRequestId` follow the identical omit-when-absent convention.
1741
+ Operational paths were never an identity input for this artifact family to
1742
+ begin with (no path parameter exists anywhere in this contract), so path
1743
+ independence holds trivially.
1744
+
1745
+ **Schema-version decision**: no bump. Every new field
1746
+ (`BoundedAgentContextSourceReferences.referenceId`/`referenceRequestId`,
1747
+ `BoundedAgentContextArtifact.fidelity`) is additive and optional; a
1748
+ pre-Prompt-7 artifact/consumer remains fully valid and behaviorally
1749
+ unchanged with all of them absent, matching the exact precedent
1750
+ `correlations?` already established without a version bump in v0.6 Batch 3.
1751
+
1752
+ **Persistence decision: none.** `projectBoundedAgentContext` and
1753
+ `projectReferenceFidelity` both remain pure, programmatic, in-memory
1754
+ functions - no new writer/reader, no new artifact family. This mirrors
1755
+ Prompt 6's own "no persisted fidelity artifact" decision and v0.6's
1756
+ existing "bounded agent context is library-only" architecture.
1757
+
1758
+ **CLI decision**: none added. v0.6 bounded agent context has never had a
1759
+ CLI surface, and this prompt does not introduce one - Prompt 8 is expected
1760
+ to become the first concrete consumer of `projectBoundedAgentContext`'s
1761
+ (now fidelity-aware) programmatic output.
1762
+
1763
+ ## v0.7 Prompt 8 controlled end-to-end external-reference coding-agent correction workflow
1764
+
1765
+ Released as `0.7.0`. One new pure domain module,
1766
+ `domain/referenceCorrectionWorkflow.ts`, plus its identity counterpart,
1767
+ `domain/referenceCorrectionIdentity.ts` - the first stage that composes
1768
+ every Prompt 1-7 and v0.1/v0.4/v0.5/v0.6 owner into one traceable
1769
+ reference-driven correction cycle. It reimplements none of them: reference
1770
+ lifecycle/adequacy (Prompt 1/3), compatibility (Prompt 4), binding (Prompt
1771
+ 5), fidelity (Prompt 6), bounded context (Prompt 7), runtime comparison
1772
+ (v0.4 `compareObservations`), and contract evaluation (v0.5
1773
+ `evaluateFrontendContract`) are all called, never re-derived. No new
1774
+ persisted artifact family, no CLI surface, no remote AI dependency, and no
1775
+ mechanism anywhere in this module (or any module it calls) that edits
1776
+ target source.
1777
+
1778
+ ```ts
1779
+ // domain/referenceCorrectionWorkflow.ts
1780
+ function prepareReferenceCorrection(input: {
1781
+ reference: ExternalReferenceArtifact; // must be approved
1782
+ baselineObservation: ObservationArtifact; // approved baseline / pre-change state
1783
+ baselineContract: PersistentBaselineContract;
1784
+ changeContract: PerChangeContract;
1785
+ bindingDeclarations: readonly ReferenceRuntimeBindingDeclaration[];
1786
+ currentObservation: ObservationArtifact; // fidelity is measured against this (= baselineObservation for the canonical proof)
1787
+ generatedAt: string; producerVersion: string; projectionProfile: ProjectionProfile;
1788
+ }): { ok: true; status: 'handoff-ready'; reviewRequestId: string; fidelity: ReferenceCandidateFidelityEvaluation; handoff: ReferenceCorrectionHandoff }
1789
+ | { ok: true; status: 'blocked-not-evaluated'; reviewRequestId: string; fidelity: ReferenceCandidateFidelityEvaluation }
1790
+ | { ok: false; reason: string };
1791
+
1792
+ function reviewReferenceCorrectionAttempt(input: {
1793
+ // ...same reference/baselineObservation/baselineContract/changeContract/bindingDeclarations...
1794
+ reviewRequestId: string; // must match the id prepareReferenceCorrection returned for this exact semantic review
1795
+ candidateObservation: ObservationArtifact; // fresh, post-edit capture
1796
+ priorAttemptId?: string;
1797
+ }): { ok: true; attempt: ReferenceCorrectionAttemptResult } | { ok: false; reason: string };
1798
+ ```
1799
+
1800
+ **Workflow architecture.** A narrowly-scoped coordinator, not a second
1801
+ workflow engine: it holds no stage catalog, no job scheduler, and no
1802
+ generic orchestration graph. It performs exactly two operations - "prepare"
1803
+ (pre-change evidence -> bounded handoff) and "review" (post-edit candidate
1804
+ -> one composed overall result) - matching this prompt's own explicit
1805
+ guidance that the external-edit boundary must remain a visible seam between
1806
+ two separate calls, never one command that blocks waiting for an external
1807
+ actor.
1808
+
1809
+ **New owners introduced**: `prepareReferenceCorrection`,
1810
+ `reviewReferenceCorrectionAttempt` (composition only - no new evaluation
1811
+ logic), `buildReferenceCorrectionReviewIdentity`/
1812
+ `buildReferenceCorrectionAttemptIdentity` (deterministic identity, see
1813
+ below), and the plain `ReferenceCorrectionHandoff`/
1814
+ `ReferenceCorrectionAttemptResult` result shapes.
1815
+
1816
+ **Existing owners reused, verbatim**: `isApprovedExternalReferenceArtifact`
1817
+ (Prompt 1), `isValidReferenceRuntimeBindingDeclarations` (Prompt 5),
1818
+ `evaluateReferenceCandidateFidelity` (Prompt 6), `projectBoundedAgentContext`
1819
+ (Prompt 7, itself now fidelity-aware), `compareObservations` (v0.4),
1820
+ `evaluateFrontendContract` (v0.5). None of their internal logic is
1821
+ inspected, duplicated, or reimplemented by this module - only their
1822
+ top-level results are read.
1823
+
1824
+ **Approved-reference/approved-baseline requirement.** `prepareReferenceCorrection`
1825
+ and `reviewReferenceCorrectionAttempt` both fail closed (`{ok: false}`) if
1826
+ `reference` is not in the `'approved'` lifecycle state (Prompt 1's own
1827
+ `isApprovedExternalReferenceArtifact` guard) - an imported-but-unapproved
1828
+ reference is never treated as an authoritative target design. Neither
1829
+ function ever calls `approveExternalReference`/`approveAndPersistBaseline`
1830
+ itself; approval remains the caller's own separate, explicit action.
1831
+
1832
+ **Pre-change candidate = approved baseline observation**, for the canonical
1833
+ proof: `prepareReferenceCorrection`'s `currentObservation` and
1834
+ `baselineObservation` are the same value, so the initial reference fidelity
1835
+ can genuinely `FAIL` (measuring the gap between the current, already-
1836
+ approved implementation and the desired new design) while the baseline
1837
+ itself stays fully valid and approved. A caller's own architecture may
1838
+ supply a distinct `currentObservation` only when justified - the workflow
1839
+ does not require them to be identical, only that `currentObservation` and
1840
+ `candidateObservation` are always independently validated
1841
+ `ObservationArtifact`s.
1842
+
1843
+ **Preparation (phase A)**: validates the common preconditions (approved
1844
+ reference, matching baseline/contract coherence, valid binding
1845
+ declarations), evaluates reference fidelity via Prompt 6 against
1846
+ `currentObservation`, and - only when that evaluation actually produced a
1847
+ result (`state !== 'not-evaluated'`) - projects it into a bounded context
1848
+ via Prompt 7/v0.6 and returns the `ReferenceCorrectionHandoff`. A
1849
+ `not-evaluated` fidelity (inadequate reference, or reference/candidate
1850
+ incompatible state) is reported as `status: 'blocked-not-evaluated'` -
1851
+ carrying the full Prompt 6 result for inspection, but never a fabricated
1852
+ handoff pretending evidence is adequate. An ambiguous or unavailable
1853
+ required binding does **not** block preparation outright - it still
1854
+ produces a `handoff-ready` result, with the ambiguity/unavailability
1855
+ visible directly in that requirement's own `unavailable`/`binding-
1856
+ unavailable` mismatch (Prompt 6's own honest per-requirement reporting,
1857
+ unchanged), so the external actor sees exactly why that specific
1858
+ requirement cannot yet be assessed.
1859
+
1860
+ **The handoff** (`ReferenceCorrectionHandoff`) carries `reviewRequestId`,
1861
+ `referenceId`/`referenceRequestId`, `baselineObservationId`,
1862
+ `currentObservationId`, the full Prompt 7 `boundedContext` (already
1863
+ containing bounded fidelity mismatches, protected/preserved context,
1864
+ adequacy/omission/truncation, and - when the caller supplied it - runtime/
1865
+ static correlation), and a fixed, four-line `verificationPlan` explaining
1866
+ in plain language what will be re-checked after the edit (fresh Chromium
1867
+ capture, reference re-evaluation, v0.4/v0.5 re-evaluation, and the exact
1868
+ overall-PASS rule) - never reduced to "make it look like the screenshot".
1869
+ No raw reference image bytes, no full `ObservationArtifact`, and no source
1870
+ excerpt are ever included.
1871
+
1872
+ **Handoff persistence: none.** The handoff is a plain, JSON-serializable,
1873
+ in-memory value returned directly to the caller - no new writer/reader, no
1874
+ new artifact family. A caller that needs the handoff to cross a process/
1875
+ session boundary (e.g. to hand it to an external coding-agent process) is
1876
+ free to serialize it with its own mechanism; observer product code does not
1877
+ own a persisted handoff artifact. This was a deliberate "smallest possible"
1878
+ choice: the handoff's only genuinely new identity is `reviewRequestId`
1879
+ (already deterministic and recomputable from stable inputs - see below), so
1880
+ nothing about it requires observer-managed persistence to remain
1881
+ traceable.
1882
+
1883
+ **Review identity** (`buildReferenceCorrectionReviewIdentity`): a pure
1884
+ function of `{referenceRequestId, baselineObservationId,
1885
+ baselineContractId, baselineContractClauses, changeContractId,
1886
+ changeContractClauses, bindingDeclarations}` only - never a timestamp,
1887
+ never an operational file path. Deliberately hashes each contract's own
1888
+ authored `clauses` content, not merely its `baselineId`/`contractId` label:
1889
+ unlike this repository's content-derived identities elsewhere (e.g.
1890
+ `ObservationArtifact.observationId`), a `PersistentBaselineContract`'s
1891
+ `baselineId` and a `PerChangeContract`'s `contractId` are plain, caller-
1892
+ authored strings (`approveAndPersistBaseline` persists `contract.baselineId`
1893
+ verbatim, never recomputing it from `clauses`) - so two structurally valid
1894
+ contracts could in principle share an id while authoring different clauses.
1895
+ Hashing clause content directly closes that gap (caught during this
1896
+ prompt's own independent-judge review before being reported PASS - see the
1897
+ report's Tooling incidents section). `reviewReferenceCorrectionAttempt`
1898
+ recomputes this same hash from its own inputs and rejects the call
1899
+ (`{ok: false}`) if the caller-supplied `reviewRequestId` does not match -
1900
+ this is the mechanism that makes "no hidden baseline change" an enforced
1901
+ invariant rather than a documented intention: an attempt claiming to belong
1902
+ to a review while actually supplying a different baseline observation,
1903
+ baseline contract (id or clause content), per-change contract (id or clause
1904
+ content), reference, or binding set can never silently succeed.
1905
+
1906
+ **Attempt identity** (`buildReferenceCorrectionAttemptIdentity`): a pure,
1907
+ deterministic function of `{reviewRequestId, candidateObservationId}` only
1908
+ - deliberately never a fresh random nonce. Every candidate observation
1909
+ already carries its own fresh, collision-resistant instance identity (v0.1's
1910
+ `buildObservationIdentity`), so hashing it together with the review it was
1911
+ captured for gives an attempt id that is both reproducible (the same
1912
+ review+candidate pair always yields the same `attemptId`) and guaranteed
1913
+ distinct per real capture.
1914
+
1915
+ **Attempt history**: caller-managed, not observer-persisted. Because both
1916
+ workflow functions are pure (no internal mutable state, no side effects),
1917
+ an already-returned `ReferenceCorrectionAttemptResult` can never be
1918
+ overwritten by a later call - a caller that keeps every attempt result it
1919
+ receives (in memory, in its own log, or in its own storage) has a complete,
1920
+ immutable, traceable history for free, linked via each attempt's own
1921
+ `reviewRequestId` (shared across all attempts of one review),
1922
+ `priorAttemptId` (an optional, purely informational link to the immediately
1923
+ preceding attempt, carried through unchanged - never consulted by the
1924
+ evaluation logic itself), and `attemptId`.
1925
+
1926
+ **Baseline-across-attempts rule**: enforced structurally, not merely
1927
+ documented. Every call to `reviewReferenceCorrectionAttempt` requires the
1928
+ caller to re-supply `baselineObservation`/`baselineContract` in full, and
1929
+ `compareObservations`/`evaluateFrontendContract` are always invoked with
1930
+ that same baseline against the fresh `candidateObservation` - there is no
1931
+ code path anywhere in this module that compares one candidate against a
1932
+ prior candidate instead. Combined with the `reviewRequestId` coherence
1933
+ check above, a caller cannot silently swap in a different baseline between
1934
+ attempts of the same review without the call being rejected.
1935
+
1936
+ **Overall result composition.** `ReferenceCorrectionOverallState =
1937
+ 'not-evaluated' | 'pass' | 'fail'`:
1938
+
1939
+ - `fidelity.state === 'not-evaluated'` -> overall `'not-evaluated'` - Prompt
1940
+ 6's own explicit blocked state is preserved exactly, never collapsed into
1941
+ an ordinary `'fail'`.
1942
+ - otherwise, `fidelity.state === 'pass' && contractEvaluation.overallVerdict === 'PASS'`
1943
+ -> overall `'pass'`; anything else -> overall `'fail'`.
1944
+
1945
+ A structurally-incomparable baseline/candidate pair is *not* given its own
1946
+ third overall bucket - v0.5's own `evaluateFrontendContract` already
1947
+ returns `'FAIL'` (never `'PASS'`) for that case, per its own established,
1948
+ unmodified precedent, and this workflow reuses that decision rather than
1949
+ re-litigating it. `approvalEligible` is a plain, read-only boolean
1950
+ (`true` iff `overallState === 'pass'`) - it is never itself an approval
1951
+ action; the caller must still invoke the existing explicit
1952
+ `approveAndPersistBaseline`/`approveExternalReference` owners separately,
1953
+ and neither is ever called from within this module.
1954
+
1955
+ **Correction iteration**: `reviewReferenceCorrectionAttempt` is called once
1956
+ per candidate; the caller decides whether and when to call it again after
1957
+ another external edit. There is no loop, no polling, no automatic retry,
1958
+ and no mechanism in this module that itself waits for or drives an external
1959
+ implementation step - the production boundary between "prepare a handoff"
1960
+ and "review a candidate" is the explicit seam a human or an external
1961
+ process controls.
1962
+
1963
+ **Source-editing boundary**: absolute. Neither this module nor anything it
1964
+ calls opens, reads, parses, or writes any target source file; both public
1965
+ operations accept only already-captured `ObservationArtifact`s and already-
1966
+ approved contract/reference artifacts. Real-Chromium candidate capture is
1967
+ always the caller's own responsibility, through the existing, unmodified
1968
+ observation pipeline (`runBrowserCapture`/`buildObservationArtifact`, the
1969
+ same functions `application/observationPersistence.ts#observe` already
1970
+ uses) - Prompt 8 adds no second browser adapter, screenshot engine, target
1971
+ resolver, or evidence-capture path.