@dailephd/my-frontend-observer 0.9.1 → 0.10.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (150) hide show
  1. package/CHANGELOG.md +490 -471
  2. package/LICENSE +21 -21
  3. package/README.md +375 -357
  4. package/dist/application/projectCheckService.d.ts +6 -0
  5. package/dist/application/projectCheckService.js +8 -1
  6. package/dist/application/projectCheckService.js.map +1 -1
  7. package/dist/application/projectWorkflowService.d.ts +7 -2
  8. package/dist/application/projectWorkflowService.js +10 -3
  9. package/dist/application/projectWorkflowService.js.map +1 -1
  10. package/dist/application/visualChangeAgentHandoffService.d.ts +28 -0
  11. package/dist/application/visualChangeAgentHandoffService.js +111 -0
  12. package/dist/application/visualChangeAgentHandoffService.js.map +1 -0
  13. package/dist/application/visualChangeProjectWorkflowService.d.ts +95 -0
  14. package/dist/application/visualChangeProjectWorkflowService.js +376 -0
  15. package/dist/application/visualChangeProjectWorkflowService.js.map +1 -0
  16. package/dist/application/visualChangeReviewService.d.ts +50 -0
  17. package/dist/application/visualChangeReviewService.js +69 -0
  18. package/dist/application/visualChangeReviewService.js.map +1 -0
  19. package/dist/application/visualChangeWorkflowPersistenceService.d.ts +26 -0
  20. package/dist/application/visualChangeWorkflowPersistenceService.js +15 -0
  21. package/dist/application/visualChangeWorkflowPersistenceService.js.map +1 -0
  22. package/dist/artifacts/visualChangeWorkflowArtifactReader.d.ts +9 -0
  23. package/dist/artifacts/visualChangeWorkflowArtifactReader.js +47 -0
  24. package/dist/artifacts/visualChangeWorkflowArtifactReader.js.map +1 -0
  25. package/dist/artifacts/visualChangeWorkflowArtifactWriter.d.ts +20 -0
  26. package/dist/artifacts/visualChangeWorkflowArtifactWriter.js +41 -0
  27. package/dist/artifacts/visualChangeWorkflowArtifactWriter.js.map +1 -0
  28. package/dist/cli.js +9 -7
  29. package/dist/cli.js.map +1 -1
  30. package/dist/domain/visualChangeAgentHandoff.d.ts +82 -0
  31. package/dist/domain/visualChangeAgentHandoff.js +80 -0
  32. package/dist/domain/visualChangeAgentHandoff.js.map +1 -0
  33. package/dist/domain/visualChangeAgentHandoffSerialization.d.ts +2 -0
  34. package/dist/domain/visualChangeAgentHandoffSerialization.js +11 -0
  35. package/dist/domain/visualChangeAgentHandoffSerialization.js.map +1 -0
  36. package/dist/domain/visualChangeCycle.d.ts +8 -0
  37. package/dist/domain/visualChangeCycle.js +7 -0
  38. package/dist/domain/visualChangeCycle.js.map +1 -0
  39. package/dist/domain/visualChangeWorkflow.d.ts +125 -0
  40. package/dist/domain/visualChangeWorkflow.js +109 -0
  41. package/dist/domain/visualChangeWorkflow.js.map +1 -0
  42. package/dist/domain/visualChangeWorkflowIdentity.d.ts +5 -0
  43. package/dist/domain/visualChangeWorkflowIdentity.js +24 -0
  44. package/dist/domain/visualChangeWorkflowIdentity.js.map +1 -0
  45. package/dist/index.d.ts +21 -1
  46. package/dist/index.js +12 -1
  47. package/dist/index.js.map +1 -1
  48. package/dist/projectWorkflow/projectPaths.d.ts +3 -0
  49. package/dist/projectWorkflow/projectPaths.js +7 -0
  50. package/dist/projectWorkflow/projectPaths.js.map +1 -1
  51. package/dist/viewer/assets/index-DglJ6f28.css +1 -0
  52. package/dist/viewer/assets/index-DsODREY5.js +9 -0
  53. package/dist/viewer/index.html +15 -15
  54. package/dist/viewer/sw.js +1 -1
  55. package/dist/viewerServer/evidence/classify.d.ts +3 -1
  56. package/dist/viewerServer/evidence/classify.js +10 -0
  57. package/dist/viewerServer/evidence/classify.js.map +1 -1
  58. package/dist/viewerServer/evidence/handles.js +1 -0
  59. package/dist/viewerServer/evidence/handles.js.map +1 -1
  60. package/dist/viewerServer/evidence/projection.d.ts +5 -0
  61. package/dist/viewerServer/evidence/projection.js +19 -0
  62. package/dist/viewerServer/evidence/projection.js.map +1 -1
  63. package/dist/viewerServer/evidence/visualChangeWorkflowView.d.ts +31 -0
  64. package/dist/viewerServer/evidence/visualChangeWorkflowView.js +36 -0
  65. package/dist/viewerServer/evidence/visualChangeWorkflowView.js.map +1 -0
  66. package/dist/viewerServer/httpServer.js +323 -1
  67. package/dist/viewerServer/httpServer.js.map +1 -1
  68. package/dist/viewerServer/referenceApproval.d.ts +22 -0
  69. package/dist/viewerServer/referenceApproval.js +42 -0
  70. package/dist/viewerServer/referenceApproval.js.map +1 -0
  71. package/dist/viewerServer/referenceVisualChangeAuthoring.d.ts +28 -0
  72. package/dist/viewerServer/referenceVisualChangeAuthoring.js +134 -0
  73. package/dist/viewerServer/referenceVisualChangeAuthoring.js.map +1 -0
  74. package/dist/viewerServer/runtimeVisualChangeAuthoring.d.ts +33 -0
  75. package/dist/viewerServer/runtimeVisualChangeAuthoring.js +81 -0
  76. package/dist/viewerServer/runtimeVisualChangeAuthoring.js.map +1 -0
  77. package/dist/viewerServer/visualChangeAuthoring.d.ts +46 -0
  78. package/dist/viewerServer/visualChangeAuthoring.js +63 -0
  79. package/dist/viewerServer/visualChangeAuthoring.js.map +1 -0
  80. package/dist/viewerServer/visualChangeHandoff.d.ts +23 -0
  81. package/dist/viewerServer/visualChangeHandoff.js +31 -0
  82. package/dist/viewerServer/visualChangeHandoff.js.map +1 -0
  83. package/dist/viewerServer/visualChangeReview.d.ts +30 -0
  84. package/dist/viewerServer/visualChangeReview.js +46 -0
  85. package/dist/viewerServer/visualChangeReview.js.map +1 -0
  86. package/docs/ARCHITECTURE.md +1394 -1373
  87. package/docs/CI_CD.md +349 -327
  88. package/docs/COMMANDS.md +1035 -1012
  89. package/docs/CONTRACTS.md +1971 -1926
  90. package/docs/CURRENT_STATE.md +1277 -1238
  91. package/docs/DEVELOPMENT.md +240 -237
  92. package/docs/DOCUMENTATION_PRESERVATION_POLICY.md +50 -50
  93. package/docs/PROJECT_DESCRIPTION.md +2248 -2224
  94. package/docs/PROJECT_MILESTONES.md +2681 -2558
  95. package/docs/PROJECT_OVERVIEW.md +200 -191
  96. package/docs/QUICKSTART.md +100 -96
  97. package/docs/RELEASE.md +37 -33
  98. package/docs/ROADMAP.md +1105 -1033
  99. package/docs/SECURITY.md +297 -275
  100. package/docs/WORKFLOWS.md +806 -770
  101. package/docs/plans/v0.10-implementation-plan.md +1509 -0
  102. package/docs/plans/v0.8-implementation-plan.md +655 -655
  103. package/docs/plans/v0.8.1-cli-usability-patch-plan.md +505 -505
  104. package/docs/plans/v0.9-implementation-plan.md +1529 -1529
  105. package/docs/plans/v0.9.1-implementation-plan.md +468 -468
  106. package/docs/reports/v0.10-batch1-visual-change-workflow-foundation.md +102 -0
  107. package/docs/reports/v0.10-batch2-project-composition-check-recording.md +103 -0
  108. package/docs/reports/v0.10-batch3-viewer-visual-change-workspace.md +93 -0
  109. package/docs/reports/v0.10-batch4-actual-frontend-entry.md +59 -0
  110. package/docs/reports/v0.10-batch5-reference-driven-entry.md +238 -0
  111. package/docs/reports/v0.10-batch6-coding-agent-handoff.md +85 -0
  112. package/docs/reports/v0.10-batch7-correction-review-acceptance.md +145 -0
  113. package/docs/reports/v0.10-batch8-integrated-acceptance.md +109 -0
  114. package/docs/reports/v0.10-implementation-completeness-documentation-reconciliation.md +344 -0
  115. package/docs/reports/v0.10-pre-release-readiness.md +120 -0
  116. package/docs/reports/v0.10-release-preparation.md +70 -0
  117. package/docs/reports/v0.10.1-project-check-baseline-context-implementation.md +86 -0
  118. package/docs/reports/v0.7-bounded-fidelity-context-prompt7.md +243 -243
  119. package/docs/reports/v0.7-implementation-completeness-documentation-reconciliation.md +497 -497
  120. package/docs/reports/v0.7-pre-release-readiness.md +337 -337
  121. package/docs/reports/v0.7-reference-binding-prompt5.md +223 -223
  122. package/docs/reports/v0.7-reference-compatibility-prompt4.md +234 -234
  123. package/docs/reports/v0.7-reference-correction-workflow-prompt8.md +222 -222
  124. package/docs/reports/v0.7-reference-fidelity-prompt6.md +216 -216
  125. package/docs/reports/v0.7-reference-foundation-prompt1.md +151 -151
  126. package/docs/reports/v0.7-reference-regions-prompt2.md +195 -195
  127. package/docs/reports/v0.7-reference-requirements-prompt3.md +217 -217
  128. package/docs/reports/v0.7-release-prep.md +423 -423
  129. package/docs/reports/v0.8-binding-fidelity-interaction-batch6.md +279 -279
  130. package/docs/reports/v0.8-bounded-context-correlation-batch7.md +233 -233
  131. package/docs/reports/v0.8-comparison-contract-inspection-batch4.md +279 -279
  132. package/docs/reports/v0.8-evidence-index-readers-batch2.md +247 -247
  133. package/docs/reports/v0.8-implementation-completeness-documentation-reconciliation.md +741 -741
  134. package/docs/reports/v0.8-integrated-viewer-acceptance-batch8.md +128 -128
  135. package/docs/reports/v0.8-observation-svg-inspection-batch3.md +223 -223
  136. package/docs/reports/v0.8-prerelease-readiness-cross-platform-security-code-rot.md +687 -687
  137. package/docs/reports/v0.8-reference-candidate-inspection-batch5.md +232 -232
  138. package/docs/reports/v0.8-viewer-runtime-pwa-batch1.md +278 -278
  139. package/docs/reports/v0.8.1-implementation-completeness-documentation-reconciliation.md +114 -114
  140. package/docs/reports/v0.8.1-prerelease-readiness-cross-platform-security-code-rot.md +170 -170
  141. package/docs/reports/v0.9-architecture-retrieval.md +14 -37
  142. package/docs/reports/v0.9-final-pre-release-readiness.md +209 -209
  143. package/docs/reports/v0.9-final-readiness-corrections.md +530 -530
  144. package/docs/reports/v0.9-pre-release-readiness.md +169 -169
  145. package/docs/reports/v0.9.1-batch1-pwa-hard-gate-isolation.md +359 -359
  146. package/docs/reports/v0.9.1-batch2-hard-gate-validation-integration.md +262 -262
  147. package/docs/reports/v0.9.1-pre-release-readiness.md +206 -206
  148. package/package.json +59 -59
  149. package/dist/viewer/assets/index-BN41MI7m.css +0 -1
  150. package/dist/viewer/assets/index-CkKXnlrI.js +0 -9
@@ -1,2224 +1,2248 @@
1
- # my-frontend-observer
2
-
3
- ## Project type
4
-
5
- Greenfield developer tool and runtime-evidence producer within the `my-dev-kit` ecosystem.
6
-
7
- ## Problem
8
-
9
- Large language models (LLMs) and coding agents can inspect frontend source code, component trees, stylesheets, project architecture, and static dependencies, but they often cannot reliably understand what an application actually looks like or how it actually behaves after a browser renders it.
10
-
11
- This creates a recurring frontend-development failure mode:
12
-
13
- 1. A user describes a visual, layout, scrolling, responsiveness, or composition problem.
14
- 2. The LLM interprets the request primarily through language and source code.
15
- 3. Static repository evidence identifies a plausible source owner.
16
- 4. A coding agent changes styling, layout, or component structure.
17
- 5. The requested local symptom appears fixed.
18
- 6. Another previously correct part of the rendered interface becomes visually or behaviorally broken.
19
- 7. Source-level tests may still pass because the regression exists only in actual browser geometry, scrolling, overflow, clipping, spacing, responsiveness, or composition.
20
- 8. The coding agent may incorrectly declare success because the requested source-level change was made without verifying the complete rendered result.
21
-
22
- Examples include:
23
-
24
- - shrinking a navigation column while leaving its contents too large for the new width;
25
- - shrinking a navigation column without transferring the released space to the intended workspace;
26
- - accidentally allowing an advertising rail to absorb released width;
27
- - changing one grid track while unintentionally moving or resizing unrelated regions;
28
- - fixing a nested scroll container in source while the rendered page still scrolls through the wrong container;
29
- - causing labels to wrap, clip, overlap, or disappear;
30
- - creating horizontal overflow at another viewport;
31
- - moving, hiding, or resizing advertising, footer, header, or workspace regions unintentionally;
32
- - satisfying one numerical styling requirement while degrading the composition as a whole;
33
- - fixing one frontend problem while silently violating a previously approved frontend behavior.
34
-
35
- Static repository understanding alone cannot reliably detect these failures because the authoritative evidence exists in the rendered browser.
36
-
37
- There is also a communication problem.
38
-
39
- A human often thinks about a frontend visually:
40
-
41
- ```text
42
- make this region narrower
43
- move this boundary
44
- give the released space to this region
45
- preserve these regions
46
- keep this scrolling behavior
47
- do not change this layout relationship
48
- ```
49
-
50
- An LLM normally receives that intent as prose and must translate it into source changes without a reliable shared representation of the rendered interface.
51
-
52
- A related failure occurs when the desired design already exists as an external visual reference. A user may have an approved PNG, WebP, screenshot, rendered mockup, or other image showing what the interface should look like, while the current application looks substantially different. Giving that image directly to a coding agent still leaves the agent to guess dimensions, spacing, region boundaries, relationships, state/theme applicability, style details, and which visual differences are actually requirements. Source/unit tests can pass while the result remains visibly far from the approved design.
53
-
54
- The system therefore needs to represent both:
55
-
56
- ```text
57
- what the browser actually rendered
58
- ```
59
-
60
- and, when supplied:
61
-
62
- ```text
63
- what an approved external visual reference specifies
64
- ```
65
-
66
- without pretending that an external raster image is an earlier runtime observation or source-code artifact.
67
-
68
- `my-frontend-observer` exists to provide that missing representation.
69
-
70
- ## Product identity
71
-
72
- `my-frontend-observer` is the runtime/browser evidence producer within the broader `my-dev-kit` ecosystem.
73
-
74
- Its responsibility is:
75
-
76
- ```text
77
- running frontend
78
- → real browser
79
- → structured runtime evidence
80
- ```
81
-
82
- It also owns the structured evidence boundary for approved external visual references used to describe desired design intent. That reference evidence remains distinct from browser runtime evidence while being comparable to a rendered candidate through explicit bindings and evaluation.
83
-
84
- It owns evidence about:
85
-
86
- - what is actually rendered;
87
- - where rendered regions are;
88
- - how large they are;
89
- - how they relate spatially;
90
- - what is visible or clipped;
91
- - what owns scrolling;
92
- - whether overflow exists;
93
- - what changed between observations;
94
- - whether approved runtime relationships remain valid;
95
- - what visual change the user intends;
96
- - what an approved external visual reference specifies where one is supplied;
97
- - which reference regions correspond to runtime targets where that binding is reliable;
98
- - how a rendered candidate differs from explicit reference-design requirements.
99
-
100
- It does not own static repository analysis.
101
-
102
- The long-term ecosystem responsibility model is:
103
-
104
- ```text
105
- my-dev-kit
106
- → static repository/source evidence producer
107
- → files
108
- → symbols
109
- → dependencies
110
- → architecture
111
- → probable source ownership
112
- → bounded source retrieval
113
-
114
- my-frontend-observer
115
- → rendered browser/runtime evidence producer
116
- → screenshots
117
- → rendered-region identity
118
- → geometry
119
- → layout relationships
120
- → scrolling and overflow
121
- → comparisons
122
- → runtime contracts
123
- → external visual-reference evidence
124
- → reference/runtime binding
125
- → visual intent
126
-
127
- my-dev-kit-orchestrator
128
- → coordinates development workflows
129
- → consumes bounded evidence when appropriate
130
- → prepares task context
131
- → manages implementation/verification workflow
132
-
133
- my-dev-kit-lab
134
- → evaluates ecosystem behavior
135
- → compatibility
136
- → controlled fixtures
137
- → experiments
138
- → evidence quality
139
- → cross-project validation
140
- ```
141
-
142
- These projects remain separately versioned and independently executable.
143
-
144
- Deep ecosystem integration must not require collapsing their responsibilities into one package.
145
-
146
- ## Product goal
147
-
148
- Build `my-frontend-observer`, a local-first frontend observation, visual-communication, and runtime-regression tool that allows humans, LLMs, coding agents, and automated checks to reason from the frontend that the browser actually rendered and, when present, from an approved external visual reference describing the intended design.
149
-
150
- The tool should convert browser state and explicitly supplied design-reference evidence into structured, inspectable evidence combining, as capabilities mature:
151
-
152
- - screenshots;
153
- - stable identities for meaningful rendered regions;
154
- - relevant Document Object Model structure;
155
- - rendered element geometry;
156
- - computed browser layout properties;
157
- - viewport information;
158
- - scroll ownership and scroll state;
159
- - visibility and overflow information;
160
- - accessibility and semantic information;
161
- - relationships between important rendered regions;
162
- - before-and-after observations;
163
- - persistent frontend invariants;
164
- - requested-change intent;
165
- - expected dependent changes;
166
- - protected regions and properties;
167
- - approved external visual references;
168
- - stable identities for meaningful reference regions;
169
- - reference geometry, relationships, applicability, tolerances, and provenance;
170
- - explicit reference-region to runtime-target bindings;
171
- - structured reference-design versus candidate evidence;
172
- - structured visual annotations.
173
-
174
- The primary goal is not to automatically redesign interfaces.
175
-
176
- The primary goal is to give a human and an LLM a shared representation of:
177
-
178
- > What is actually on the screen, where it is, how large it is, how it behaves, how its important regions relate to one another, what the user wants changed or wants matched from an approved reference, what must remain intact, and what actually changed after an implementation edit?
179
-
180
- ## Three primary product jobs
181
-
182
- ### 1. Human-to-LLM design and layout communication
183
-
184
- The observer should help a user communicate visual and layout intent without requiring the LLM to infer everything from prose or source code.
185
-
186
- The system should eventually allow communication through a combination of:
187
-
188
- ```text
189
- runtime screenshot or approved external reference
190
- + stable named regions
191
- + measured or authored geometry
192
- + layout relationships
193
- + runtime behavior where applicable
194
- + visual annotation
195
- + textual intent
196
- ```
197
-
198
- A user should be able to communicate ideas such as:
199
-
200
- ```text
201
- make primary navigation narrower
202
- give the released horizontal space to the workspace
203
- preserve both advertising rails
204
- do not clip navigation labels
205
- keep document-level scrolling
206
- leave the footer relationship unchanged
207
- ```
208
-
209
- or:
210
-
211
- ```text
212
- make this popup match the approved One Dark reference
213
- match this card's width and horizontal position
214
- preserve these button proportions and spacing
215
- use this approved artwork rather than redesigning it
216
- this footer placement is informational, not an exact requirement
217
- ```
218
-
219
- without needing to express the implementation mechanism.
220
-
221
- An external reference is desired-design evidence, not implementation. A raster image does not reveal original DOM structure, CSS, vector paths, component hierarchy, source ownership, hidden layout constraints, or design-token names unless those are separately supplied. The observer must preserve that distinction rather than fabricating hidden structure.
222
-
223
- ### 2. Safe LLM-assisted frontend changes
224
-
225
- The observer should make the complete rendered result part of the definition of implementation success.
226
-
227
- A requested local change must not be considered successful merely because the requested element changed.
228
-
229
- The tool should eventually distinguish:
230
-
231
- ```text
232
- requested change
233
- expected dependent change
234
- protected change
235
- preserved invariant
236
- unexpected change
237
- ```
238
-
239
- This creates an explicit allowed scope of frontend change.
240
-
241
- Previously approved runtime behavior must remain valid unless the user explicitly supersedes it.
242
-
243
- An external reference-fidelity success must obey the same rule. Matching a supplied design region does not authorize breaking a protected control, preserved relationship, baseline invariant, or unrelated application behavior.
244
-
245
- ### 3. Runtime evidence for the my-dev-kit ecosystem
246
-
247
- The observer should provide the runtime evidence domain that static repository analysis cannot provide.
248
-
249
- The intended long-term combination is:
250
-
251
- ```text
252
- my-dev-kit static evidence
253
- +
254
- my-frontend-observer runtime evidence
255
- +
256
- my-frontend-observer reference evidence where applicable
257
- ↓
258
- bounded coordinated context
259
- ↓
260
- LLM / coding agent
261
- ```
262
-
263
- A runtime region may eventually be correlated with bounded source evidence without requiring the observer to become a source-analysis engine or requiring `my-dev-kit` to become a browser runner. Likewise, a reference region may bind to a runtime target without implying source ownership.
264
-
265
- ## Intended users
266
-
267
- Primary users:
268
-
269
- - developers using LLMs or coding agents for frontend development;
270
- - developers debugging visual and responsive regressions;
271
- - developers who need to communicate visual layout intent to an LLM;
272
- - developers who need to reproduce or validate an approved external visual design reference;
273
- - developers who need browser-observed evidence before accepting frontend changes;
274
- - maintainers who want machine-readable runtime frontend evidence;
275
- - maintainers of the broader `my-dev-kit` ecosystem.
276
-
277
- The initial user is a developer working locally with web applications, LLMs, and coding agents.
278
-
279
- ## Initial workflow
280
-
281
- The first useful workflow is intentionally small:
282
-
283
- ```text
284
- local web application
285
- → target URL + viewport + explicit observation targets
286
- → my-frontend-observer launches Chromium
287
- → browser renders target
288
- → observer captures screenshot
289
- → observer captures structured page evidence
290
- → observer captures structured target evidence
291
- → observer writes a versioned local observation artifact
292
- → command-line result reports completion, warnings, or failure
293
- ```
294
-
295
- The first version does not need comparison, change contracts, visual annotation, source ownership, external visual-reference evaluation, or direct integration with other ecosystem projects.
296
-
297
- Its purpose is to establish a trustworthy runtime-evidence foundation.
298
-
299
- ## Intended end-to-end workflow
300
-
301
- The long-term workflow supports two human intent entry modes.
302
-
303
- Actual-frontend-driven:
304
-
305
- ```text
306
- local web application
307
- ↓
308
- my-frontend-observer
309
- ↓
310
- baseline observation
311
- + screenshot
312
- + regions
313
- + geometry
314
- + relationships
315
- + runtime behavior
316
- ↓
317
- human reviews frontend
318
- ↓
319
- human requests or visually annotates change
320
- ```
321
-
322
- Reference-driven:
323
-
324
- ```text
325
- approved external visual reference
326
- + current browser-rendered candidate
327
- ↓
328
- reference identity, regions, relationships, applicability, and intent
329
- ↓
330
- explicit reference-region ↔ runtime-target binding
331
- ↓
332
- structured reference-vs-candidate evidence
333
- ↓
334
- human confirms which reference details are requirements
335
- ```
336
-
337
- Both then converge on:
338
-
339
- ```text
340
- requested / dependent / protected / preserved change scope
341
- ↓
342
- bounded runtime evidence
343
- + relevant reference evidence where applicable
344
- +
345
- bounded static evidence from my-dev-kit where useful
346
- ↓
347
- my-dev-kit-orchestrator / developer / LLM
348
- ↓
349
- coding agent changes target source separately
350
- ↓
351
- my-frontend-observer captures new state
352
- ↓
353
- before/after comparison
354
- + reference-vs-candidate evaluation where applicable
355
- ↓
356
- requested changes evaluated
357
- + dependent changes evaluated
358
- + protected properties evaluated
359
- + existing regression contracts rerun
360
- + reference requirements evaluated
361
- ↓
362
- PASS
363
- or
364
- actionable evidence identifying what broke or still differs
365
- ↓
366
- human approves new baseline/reference state
367
- or requests another iteration
368
- ```
369
-
370
- The target application remains a separate project throughout this process.
371
-
372
- ## Principal capability 1 — Browser observation
373
-
374
- The tool must observe a locally running web application through a real browser.
375
-
376
- Initial browser support should use Chromium through Playwright unless architecture work establishes a materially better supported mechanism.
377
-
378
- The initial implementation should accept at minimum:
379
-
380
- - target URL;
381
- - viewport width;
382
- - viewport height;
383
- - explicitly configured observation targets;
384
- - output location.
385
-
386
- Later configuration may support:
387
-
388
- - route collections;
389
- - themes;
390
- - reusable scenarios;
391
- - browser-state setup;
392
- - authentication setup;
393
- - device profiles;
394
- - interaction sequences.
395
-
396
- Those later capabilities are not required for the first version.
397
-
398
- Browser runtime behavior is authoritative for rendered geometry.
399
-
400
- The observer must not infer final layout solely from source styles.
401
-
402
- ## Principal capability 2 — Screenshot capture
403
-
404
- For each observation, capture the rendered page as an image.
405
-
406
- The screenshot is evidence associated with the same observation identity as the structured browser measurements.
407
-
408
- Screenshots support:
409
-
410
- - human review;
411
- - multimodal LLM review;
412
- - annotation;
413
- - before/after inspection;
414
- - reference/candidate inspection;
415
- - regression evidence.
416
-
417
- The system should eventually support:
418
-
419
- - viewport screenshots;
420
- - full-page screenshots where useful.
421
-
422
- Exact initial screenshot behavior and capture-readiness semantics must be defined before implementation.
423
-
424
- Pixel-perfect screenshot comparison must not become the only regression mechanism or the only reference-fidelity mechanism.
425
-
426
- Structured browser evidence remains essential.
427
-
428
- ## Principal capability 3 — Stable rendered-region identity
429
-
430
- Meaningful rendered regions need stable logical identities so humans, LLMs, comparisons, annotations, regression contracts, and reference bindings can refer to the same conceptual runtime region over time.
431
-
432
- Examples may include:
433
-
434
- ```text
435
- app-shell
436
- header
437
- primary-navigation
438
- main-content
439
- tool-workspace
440
- left-ad-rail
441
- right-ad-rail
442
- footer-ad
443
- footer
444
- theme-control
445
- ```
446
-
447
- These names are examples only.
448
-
449
- The observer must not assume that every application uses the same regions.
450
-
451
- Region identity should support appropriate browser-observable mechanisms such as:
452
-
453
- - semantic HTML elements;
454
- - accessibility role;
455
- - accessible name;
456
- - stable `id`;
457
- - stable `data-*` attribute;
458
- - bounded CSS selector fallback;
459
- - text-based selection only where appropriate.
460
-
461
- A target may have a stable observer-level identity without having a known source-code component identity.
462
-
463
- For example:
464
-
465
- ```text
466
- runtime target:
467
- primary-navigation
468
- ```
469
-
470
- does not by itself prove:
471
-
472
- ```text
473
- source owner:
474
- VerticalNav.tsx
475
- ```
476
-
477
- Source ownership belongs to the static-analysis integration boundary.
478
-
479
- A reference region likewise has a separate reference identity. The released
480
- v0.7 binding model keeps:
481
-
482
- ```text
483
- reference region:
484
- primary-navigation-reference
485
-
486
- runtime target:
487
- primary-navigation
488
- ```
489
-
490
- as two explicit identity domains rather than collapsing them into one.
491
-
492
- ## Principal capability 4 — Rendered layout map
493
-
494
- Capture a structured representation of important rendered elements.
495
-
496
- For an observed region, useful browser evidence includes:
497
-
498
- ```text
499
- identifier
500
- selection method
501
- semantic role
502
- tag
503
- accessible name where available
504
- text summary where appropriate
505
-
506
- x
507
- y
508
- width
509
- height
510
- right
511
- bottom
512
-
513
- visibility
514
- display
515
- position
516
- overflow-x
517
- overflow-y
518
- z-index where relevant
519
-
520
- scroll width
521
- scroll height
522
- client width
523
- client height
524
- scroll top
525
- scroll left
526
- ```
527
-
528
- The observer should prefer browser-computed values over attempting to infer final geometry from source styling.
529
-
530
- Observed dimensions are measurements, not automatically design constants.
531
-
532
- For example:
533
-
534
- ```text
535
- primary-navigation.width = 176
536
- ```
537
-
538
- means:
539
-
540
- ```text
541
- the browser rendered the observed region at 176 pixels
542
- ```
543
-
544
- It does not automatically mean:
545
-
546
- ```text
547
- navigation must always be exactly 176 pixels wide
548
- ```
549
-
550
- Responsive layouts must remain possible.
551
-
552
- The output must distinguish:
553
-
554
- ```text
555
- direct browser observation
556
- computed browser property
557
- derived relationship or interpretation
558
- ```
559
-
560
- ## Principal capability 5 — Page-level browser state
561
-
562
- Capture page-level evidence such as:
563
-
564
- ```text
565
- URL
566
- final URL after navigation
567
- document title
568
-
569
- viewport width
570
- viewport height
571
- device pixel ratio
572
-
573
- document width
574
- document height
575
- document scroll width
576
- document scroll height
577
- document client width
578
- document client height
579
-
580
- window scroll X
581
- window scroll Y
582
-
583
- horizontal overflow state
584
- vertical overflow state
585
- ```
586
-
587
- This should make questions such as these answerable from runtime evidence:
588
-
589
- - Does the document own vertical scrolling?
590
- - Is a child container actually scrolling instead?
591
- - Is there horizontal document overflow?
592
- - Is the footer below the initial viewport?
593
- - Did the page become taller or wider after a change?
594
- - Did viewport behavior change unexpectedly?
595
-
596
- Reference-driven evaluation uses explicit caller-supplied state/applicability
597
- identity so the observer does not compare the wrong theme, viewport,
598
- authentication state, or application state as though it were the intended
599
- reference state.
600
-
601
- ## Principal capability 6 — Runtime scrolling, overflow, and visibility
602
-
603
- Scrolling must be treated as runtime behavior rather than inferred solely from style declarations.
604
-
605
- The observer should eventually be able to:
606
-
607
- 1. capture initial scroll state;
608
- 2. perform a controlled scroll action;
609
- 3. capture resulting scroll state;
610
- 4. identify which observed regions changed scroll position;
611
- 5. expose evidence about which container appears to own scrolling;
612
- 6. identify whether elements enter or leave the viewport;
613
- 7. identify horizontal or vertical overflow.
614
-
615
- Example direct observation:
616
-
617
- ```text
618
- before:
619
- window.scrollY = 0
620
- main.scrollTop = 0
621
-
622
- after requested page scroll:
623
- window.scrollY = 500
624
- main.scrollTop = 0
625
- ```
626
-
627
- Possible derived interpretation:
628
-
629
- ```text
630
- document appears to own primary vertical scrolling
631
- ```
632
-
633
- The observer must not present the derived statement as if it were a direct browser measurement.
634
-
635
- ## Principal capability 7 — Layout relationships and dependency relationships
636
-
637
- Individual measurements are not enough.
638
-
639
- Many design requirements concern relationships between regions.
640
-
641
- The observer should support relationship-oriented evidence such as:
642
-
643
- ```text
644
- navigation is left of workspace
645
- workspace is wider than navigation
646
- navigation does not overlap workspace
647
- workspace does not overlap right advertising rail
648
- footer begins after main content
649
- element is contained inside parent
650
- navigation contents fit inside navigation
651
- document width does not exceed viewport width
652
- ```
653
-
654
- The system should also leave room for an explicit layout relationship or dependency model.
655
-
656
- Example:
657
-
658
- ```text
659
- Viewport
660
- ↓
661
- AppShell
662
- ├── LeftAd
663
- ├── Navigation
664
- ├── Workspace
665
- └── RightAd
666
- ```
667
-
668
- A requested change may imply legitimate dependent changes.
669
-
670
- Example:
671
-
672
- ```text
673
- Navigation width decreases
674
- ↓
675
- Workspace width increases
676
- Workspace x-position may move
677
- ```
678
-
679
- Other properties may need to remain preserved:
680
-
681
- ```text
682
- LeftAd width
683
- RightAd width
684
- Header height
685
- Footer relationships
686
- ```
687
-
688
- The system must distinguish observed relationships from causal claims.
689
-
690
- It should not automatically claim that one region caused another region to change merely because both changed.
691
-
692
- Expected dependency semantics should come from an explicit contract, user intent, approved reference intent, or another supported source of evidence.
693
-
694
- Reference regions reuse the canonical relationship vocabulary when the same
695
- geometric relation applies, while preserving the fact that reference
696
- relationships are derived from explicit reference-image geometry rather than
697
- browser-observed DOM/runtime facts.
698
-
699
- ## Principal capability 8 — Observation artifact
700
-
701
- Each capture should produce one cohesive, observer-owned, versioned observation artifact or artifact directory.
702
-
703
- The exact schema and filenames must be decided during architecture and schema design.
704
-
705
- A conceptual structure may resemble:
706
-
707
- ```text
708
- observation/
709
- manifest.json
710
- page.json
711
- elements.json
712
- screenshot.png
713
- ```
714
-
715
- Possible future additions may include:
716
-
717
- ```text
718
- relationships.json
719
- interactions.json
720
- comparison.json
721
- contracts.json
722
- annotations.json
723
- summary.txt
724
- ```
725
-
726
- These names are conceptual rather than fixed requirements.
727
-
728
- The public artifact contract should establish from the beginning:
729
-
730
- ```text
731
- artifact kind
732
- schema version
733
- observation identity
734
- producer version
735
- browser identity
736
- request/configuration identity
737
- provenance
738
- artifact references
739
- completion state
740
- diagnostics
741
- limits
742
- truncation/omission reporting
743
- ```
744
-
745
- Artifact paths should be relative and portable where possible.
746
-
747
- Heavy evidence such as screenshots should be referenced rather than embedded into unrelated structured records.
748
-
749
- Consumers must be able to distinguish a completed observation from a partial or failed capture.
750
-
751
- The artifact must distinguish:
752
-
753
- ```text
754
- observed evidence
755
- derived evidence
756
- unavailable evidence
757
- not-applicable evidence
758
- partial evidence
759
- ```
760
-
761
- The artifact schema should evolve intentionally and additively where compatible.
762
-
763
- Package version and observation schema version must remain separate concepts.
764
-
765
- The v0.7 `ExternalReferenceArtifact` is a separate evidence family. It does
766
- not masquerade as an observation merely to reuse an existing serializer.
767
-
768
- ## Principal capability 9 — Before/after comparison
769
-
770
- The tool should compare two observations representing comparable logical frontend states.
771
-
772
- Useful differences include:
773
-
774
- ```text
775
- element moved
776
- element resized
777
- element disappeared
778
- element appeared
779
- visibility changed
780
- element became clipped
781
- horizontal overflow appeared
782
- vertical overflow changed
783
- document size changed
784
- scroll-owner evidence changed
785
- relative position changed
786
- layout relationship changed
787
- ```
788
-
789
- Comparison should produce structured evidence such as:
790
-
791
- ```text
792
- target
793
- property or relationship
794
- before value
795
- after value
796
- difference
797
- classification
798
- supporting observation identities
799
- ```
800
-
801
- Example:
802
-
803
- ```text
804
- Target: primary-navigation
805
- Property: width
806
- Before: 176
807
- After: 97
808
- Difference: -79
809
- ```
810
-
811
- The comparison engine should preserve references to before/after screenshots and underlying observations.
812
-
813
- The tool should not rely solely on screenshot pixel differences.
814
-
815
- Before/after comparison is not part of the first observation version.
816
-
817
- The initial observation identity and provenance model must nevertheless preserve enough information to support future comparability decisions.
818
-
819
- Before/after comparison remains conceptually distinct from reference-design
820
- versus candidate evaluation. An external desired-state image is not an earlier
821
- runtime state.
822
-
823
- ## Principal capability 10 — Explicit change scope
824
-
825
- A central long-term concept is the ability to represent what a requested frontend change is allowed to affect.
826
-
827
- A change should be expressible through categories such as:
828
-
829
- ### Requested changes
830
-
831
- Properties or relationships explicitly intended to change.
832
-
833
- Example:
834
-
835
- ```text
836
- primary-navigation.width
837
- → decrease significantly
838
- ```
839
-
840
- ### Expected dependent changes
841
-
842
- Properties expected to change as a legitimate consequence.
843
-
844
- Example:
845
-
846
- ```text
847
- tool-workspace.width
848
- → increase using released horizontal space
849
-
850
- tool-workspace.x
851
- → may move left
852
- ```
853
-
854
- ### Protected properties or regions
855
-
856
- Properties expected to remain unchanged.
857
-
858
- Example:
859
-
860
- ```text
861
- left-ad-rail.width
862
- right-ad-rail.width
863
- header.height
864
- ```
865
-
866
- ### Preserved invariants and behaviors
867
-
868
- Previously correct relationships or behaviors that must remain true.
869
-
870
- Example:
871
-
872
- ```text
873
- navigation contents remain unclipped
874
- navigation does not overlap workspace
875
- workspace does not overlap advertising rails
876
- document does not horizontally overflow
877
- document continues to own primary page scrolling
878
- mobile layout remains usable
879
- ```
880
-
881
- Together, these categories define the allowed scope of rendered change.
882
-
883
- This concept may eventually be represented by an explicit `ChangeContract` or equivalent schema.
884
-
885
- The conceptual name does not require that exact implementation type.
886
-
887
- Reference-derived executable intent maps into this same change-scope model. A
888
- visible detail in a reference may remain informational or unassessed until
889
- explicitly promoted into requested, expected-dependent, protected, or
890
- preserved intent. There is no separate reference-only change taxonomy.
891
-
892
- ## Principal capability 11 — Frontend regression and change contracts
893
-
894
- The project should support persistent executable runtime invariants.
895
-
896
- Examples include:
897
-
898
- ```text
899
- element is visible
900
- element is not clipped
901
- element width is within a bound
902
- element A does not overlap element B
903
- element A is wider than element B
904
- element A follows element B vertically
905
- document width does not exceed viewport width
906
- window owns requested page scrolling
907
- specified element does not own primary page scrolling
908
- element begins below initial viewport
909
- ```
910
-
911
- Relationship-oriented contracts should be preferred when they represent user intent more accurately than fixed pixels.
912
-
913
- For example:
914
-
915
- Prefer:
916
-
917
- ```text
918
- workspace width increases when navigation width decreases
919
- ```
920
-
921
- when that is the actual design requirement.
922
-
923
- Use:
924
-
925
- ```text
926
- navigation.width = 97
927
- ```
928
-
929
- only when the user truly requires that exact value.
930
-
931
- The system should support two related forms of contract:
932
-
933
- ```text
934
- persistent baseline contracts
935
- ```
936
-
937
- and:
938
-
939
- ```text
940
- per-change contracts
941
- ```
942
-
943
- Persistent baseline contracts preserve approved frontend behavior across future changes.
944
-
945
- Per-change contracts describe:
946
-
947
- ```text
948
- requested changes
949
- expected dependent changes
950
- protected properties
951
- preserved invariants
952
- ```
953
-
954
- Example evaluation:
955
-
956
- ```text
957
- REQUESTED CHANGE
958
- Navigation.width
959
- 176 → 97
960
- PASS
961
-
962
- EXPECTED DEPENDENT CHANGE
963
- Workspace.width
964
- 960 → 1039
965
- PASS
966
-
967
- PROTECTED PROPERTY
968
- RightAd.width
969
- 112 → 154
970
- FAIL
971
-
972
- PRESERVED INVARIANT
973
- Navigation content became clipped
974
- FAIL
975
-
976
- OVERALL
977
- FAIL
978
- ```
979
-
980
- A frontend change must not be declared successful merely because its requested local mutation succeeded.
981
-
982
- Reference fidelity supplements these contracts. It does not replace or weaken
983
- them, and a fidelity pass cannot override a protected or preserved contract
984
- failure.
985
-
986
- ## Principal capability 12 — Bounded agent context and static/runtime integration
987
-
988
- Structured output must support both programmatic use and LLM consumption.
989
-
990
- The tool should eventually produce a bounded runtime-evidence package containing, as applicable:
991
-
992
- - target page identity;
993
- - viewport;
994
- - major observed regions;
995
- - region geometry;
996
- - semantic identities;
997
- - layout relationships;
998
- - dependency/change-scope information;
999
- - overflow state;
1000
- - scroll evidence;
1001
- - comparison results;
1002
- - contract results;
1003
- - important warnings;
1004
- - references to underlying raw evidence.
1005
-
1006
- Preserve the evidence hierarchy:
1007
-
1008
- ```text
1009
- raw browser evidence
1010
- ↓
1011
- normalized structured evidence
1012
- ↓
1013
- derived relationships
1014
- ↓
1015
- bounded summary/context
1016
- ↓
1017
- LLM reasoning
1018
- ```
1019
-
1020
- The bounded context must not require an LLM to consume:
1021
-
1022
- - an entire raw Document Object Model dump;
1023
- - every computed style property;
1024
- - enormous accessibility trees;
1025
- - repeated unchanged measurements;
1026
- - every screenshot produced during a workflow.
1027
-
1028
- The summary must remain traceable to the evidence supporting it.
1029
-
1030
- The initial command-line version may return a concise execution summary.
1031
-
1032
- That operational summary must not be confused with the richer agent-oriented context package described here.
1033
-
1034
- The shortest path to practical coding-agent use combines this bounded runtime
1035
- projection with relevant bounded static/source evidence from `my-dev-kit`.
1036
- The evidence domains remain separate and traceable:
1037
-
1038
- ```text
1039
- observer runtime evidence
1040
- +
1041
- my-dev-kit static evidence
1042
- ↓
1043
- bounded agent context
1044
- ↓
1045
- external coding agent
1046
- ```
1047
-
1048
- When an external reference is active, the released v0.7 extension adds only
1049
- task-relevant reference identity, selected design requirements, measurable
1050
- candidate mismatches, bound runtime targets, protected/preserved context, and
1051
- references to heavy image assets. It does not place the full reference artifact
1052
- or every image difference into the agent packet by default.
1053
-
1054
- Runtime/static correlation must be explicit and may be ambiguous. A stable
1055
- runtime target identity must never silently become a source-ownership claim.
1056
- The observer owns runtime projection and its correlation/export boundary;
1057
- `my-dev-kit` owns static indexing and retrieval; the orchestrator coordinates
1058
- bounded consumption; the lab owns exact compatibility evaluation.
1059
-
1060
- This integrated, text/config-driven path supports an end-to-end coding-agent
1061
- change review before the viewer or visual annotation becomes a prerequisite.
1062
- The observer does not edit source: an external coding agent makes the change,
1063
- after which the observer rerenders, compares, and evaluates preserved contracts
1064
- and, where applicable, reference fidelity.
1065
-
1066
- ## Implemented capability (v0.7) — External visual reference and reference-driven design evidence
1067
-
1068
- v0.7 supports approved external visual references as a structured desired-design evidence domain.
1069
-
1070
- The public `import-reference` command accepts PNG, JPEG, and WebP images. Format
1071
- and dimensions are detected from bounded header bytes rather than trusted from a
1072
- filename extension, and the implementation enforces bounded file-size and image-
1073
- dimension limits. No OCR, image segmentation, computer-vision target discovery,
1074
- or raster-to-code reconstruction is part of this capability.
1075
-
1076
- A raw reference image is evidence, not implementation. It does not reveal hidden DOM structure, source ownership, original CSS, component hierarchy, design tokens, original vector paths, or inaccessible font metadata.
1077
-
1078
- The `ExternalReferenceArtifact` and its derived evaluation path preserve enough
1079
- structured information to answer:
1080
-
1081
- ```text
1082
- which exact reference image/version was used?
1083
- which regions were defined?
1084
- which requirements were explicitly authored?
1085
- which relationships were derived?
1086
- which candidate observation was evaluated?
1087
- which viewport/theme/application state applies?
1088
- which tolerance/evaluation policy was used?
1089
- which approval or supersession decision applies?
1090
- ```
1091
-
1092
- ### Reference-region model
1093
-
1094
- v0.7 reference regions are explicit, bounded semantic rectangles authored by a
1095
- user or configuration. Each region has a stable `id` and a canonical
1096
- `{x, y, width, height}` rectangle in reference-image pixels with origin at the
1097
- image's top-left corner. `right`, `bottom`, `centerX`, and `centerY` are derived
1098
- on demand from that canonical rectangle and are not redundantly persisted.
1099
- There is no automatic segmentation and no normalized-coordinate region model in
1100
- v0.7.
1101
-
1102
- Reference regions reuse the same geometry-only relationship predicates used by
1103
- runtime layout relationships where the concept is genuinely shared, including
1104
- horizontal/vertical order, overlap, relative width, geometric fit, and vertical
1105
- sequencing. Reference relationships are derived on demand and do not become
1106
- requirements automatically.
1107
-
1108
- ### Reference-design intent and tolerances
1109
-
1110
- Not every visible pixel is a requirement.
1111
-
1112
- Reference evidence distinguishes:
1113
-
1114
- ```text
1115
- visible/derived evidence
1116
- explicit authored requirement
1117
- informational/unassessed detail
1118
- ```
1119
-
1120
- Executable reference requirements reuse the canonical v0.5 authored categories:
1121
-
1122
- ```text
1123
- requested
1124
- expected-dependent
1125
- protected
1126
- preserved
1127
- ```
1128
-
1129
- `unexpected` remains derived-only.
1130
-
1131
- v0.7 supports selected requirement subjects over region properties,
1132
- region-to-region relationships, and bounded two-region measurements. Numeric
1133
- reference tolerances are explicitly reference-owned and use:
1134
-
1135
- - `exact`;
1136
- - `absolute-reference-px`;
1137
- - `percent`.
1138
-
1139
- Reference-image coordinates and tolerances are not silently treated as CSS
1140
- pixels. Fidelity establishes an explicit full-frame reference-image-pixel to
1141
- CSS-pixel scale from the reference image dimensions and declared applicable
1142
- runtime viewport, with an independent aspect-ratio-coherence gate.
1143
-
1144
- Selected color/style evidence, asset-similarity evidence, or image-region
1145
- similarity are not v0.7 success mechanisms. They may be added later only as
1146
- bounded supplemental evidence and must not replace structured geometry,
1147
- relationships, applicability, or canonical contract evaluation.
1148
-
1149
- ### Reference applicability and comparability
1150
-
1151
- Reference and candidate must represent compatible intended states before ordinary fidelity differences are evaluated.
1152
-
1153
- v0.7 supports explicit caller-supplied applicability dimensions for:
1154
-
1155
- ```text
1156
- viewport
1157
- theme
1158
- application state
1159
- authenticated state
1160
- ```
1161
-
1162
- The corresponding candidate state is likewise caller/configuration supplied on
1163
- observation. It is not inferred from screenshot pixels, DOM, CSS, URL, or source
1164
- code.
1165
-
1166
- For example:
1167
-
1168
- ```text
1169
- reference: One Dark / active crawl
1170
- candidate: One Light / idle
1171
- ```
1172
-
1173
- produces an explicit incompatible/incomparable result rather than a meaningless visual-difference list when those dimensions are declared and conflict.
1174
-
1175
- The compatibility implementation reuses the existing v0.4 comparability result
1176
- and per-dimension comparison conventions rather than creating an unrelated
1177
- reference-only state system.
1178
-
1179
- ### Reference-to-runtime binding
1180
-
1181
- The system uses an explicit association between:
1182
-
1183
- ```text
1184
- reference region
1185
- ```
1186
-
1187
- and:
1188
-
1189
- ```text
1190
- runtime target
1191
- ```
1192
-
1193
- Binding declarations are caller/configuration supplied and are never inferred
1194
- from geometry, matching names, or source code. Binding results use the closed
1195
- states:
1196
-
1197
- ```text
1198
- bound
1199
- ambiguous
1200
- unavailable
1201
- ```
1202
-
1203
- Reference identity, runtime identity, and source identity remain separate domains.
1204
-
1205
- ### Structured reference-vs-candidate evaluation
1206
-
1207
- v0.7 combines selected reference requirements with browser-authoritative
1208
- candidate evidence and produces bounded, actionable structured fidelity results.
1209
- Evaluation proceeds through reference structural validation, reference-evidence
1210
- adequacy, reference/candidate compatibility, explicit binding, then each
1211
- selected requirement. An inadequate reference or incompatible candidate is
1212
- `not-evaluated`; it is not fabricated into an ordinary visual failure.
1213
-
1214
- Example evidence remains of the form:
1215
-
1216
- ```text
1217
- Target: current-page-card
1218
- Reference x: 28
1219
- Candidate x: 18
1220
- Delta: -10
1221
-
1222
- Reference width: 424
1223
- Candidate width: 446
1224
- Delta: +22
1225
-
1226
- Expected separation below header: 24px within tolerance
1227
- Candidate separation: 38px
1228
- Result: fidelity requirement failed
1229
- ```
1230
-
1231
- Pixel or image similarity is not the success mechanism in v0.7.
1232
-
1233
- ### Reference lifecycle and approval
1234
-
1235
- A random supplied image never silently becomes a project baseline or active design authority.
1236
-
1237
- The v0.7 persisted lifecycle has exactly two explicit states:
1238
-
1239
- ```text
1240
- imported
1241
- → approved
1242
- ```
1243
-
1244
- `import-reference` creates a new imported artifact. `approve-reference` is the
1245
- only explicit approval act and creates a new approved artifact instance while
1246
- preserving the imported artifact unchanged. Supersession is represented by a
1247
- forward pointer on the newer artifact and never rewrites the superseded
1248
- artifact. There is no automatic measured, annotated, active, or auto-approved
1249
- lifecycle state in v0.7.
1250
-
1251
- Reference approval is separate from baseline approval. Reference supersession is separate from baseline supersession. A fidelity `PASS` does not approve either one automatically.
1252
-
1253
- ### Multiple references
1254
-
1255
- The model permits separately identified references for explicit states such as:
1256
-
1257
- ```text
1258
- dark theme / idle
1259
- dark theme / active
1260
- dark theme / error
1261
- light theme / idle
1262
- light theme / active
1263
- light theme / error
1264
- desktop
1265
- mobile
1266
- ```
1267
-
1268
- The correct reference must be selected by explicit identity/applicability rules rather than by accidental filename matching.
1269
-
1270
- ### Asset fidelity
1271
-
1272
- A reference region may represent artwork or another asset-sensitive area.
1273
-
1274
- The observer can currently preserve the reference image and structured region
1275
- geometry but does not claim to recover vector paths or hidden source data from a
1276
- raster reference. Raster-to-vector reconstruction and image-to-code generation
1277
- remain external implementation concerns. Bounded asset/image-similarity evidence
1278
- would be a later extension, not a current v0.7 contract.
1279
-
1280
- ### Coding-agent correction packet
1281
-
1282
- The v0.7 bounded agent context reports measurable reference/candidate mismatches
1283
- rather than asking the coding agent to reinterpret the entire image each
1284
- iteration. It carries relevant failed requirements, bound runtime targets,
1285
- active protected/preserved context, provenance, adequacy/omission/truncation,
1286
- and bounded static/source correlation when supplied by the caller.
1287
-
1288
- The implemented correction loop is:
1289
-
1290
- ```text
1291
- approved external reference
1292
- → structured reference evidence
1293
- → current candidate observation
1294
- → structured fidelity mismatch
1295
- → bounded runtime/static context
1296
- → external coding agent correction
1297
- → real Chromium rerender
1298
- → reevaluate reference fidelity
1299
- + canonical before/after comparison
1300
- + canonical baseline/per-change contract evaluation
1301
- → PASS or actionable failure
1302
- ```
1303
-
1304
- Matching the reference is necessary but never sufficient: an active protected or
1305
- preserved contract regression still makes the overall correction review fail.
1306
- The observer never edits target source; the implementation actor remains external.
1307
-
1308
- The viewer and annotation systems later consume this reference model. They must not create another one.
1309
-
1310
- ## Implemented foundation (v0.6) — Static/runtime source association
1311
-
1312
- The observer provides an explicit programmatic runtime/static correlation
1313
- boundary for associating stable runtime targets with caller-supplied bounded
1314
- static candidates where reliable.
1315
-
1316
- The chain is:
1317
-
1318
- ```text
1319
- rendered region
1320
- → runtime target identity
1321
- → correlation evidence
1322
- → my-dev-kit static identity / bounded evidence
1323
- → relevant source retrieval
1324
- ```
1325
-
1326
- Correlation results preserve `correlated`, `ambiguous`, or `unavailable`
1327
- outcomes and competing candidates. The observer does not implement a competing
1328
- repository-analysis system and does not silently turn a runtime target into a
1329
- source owner. `my-dev-kit` remains the owner of repository crawling, parsing,
1330
- indexing, source graphs, architecture, and bounded retrieval. The observer
1331
- package has no runtime dependency on `@dailephd/my-dev-kit`; static candidate
1332
- evidence is supplied through the explicit boundary.
1333
-
1334
- ## Principal capability 13 — Human visual review
1335
-
1336
- After the text/config-driven coding-agent workflow and non-graphical external-reference evidence foundation are proven, the project should provide a human-readable graphical way to inspect the same canonical evidence.
1337
-
1338
- A later local interface should allow the developer to:
1339
-
1340
- - view the captured screenshot;
1341
- - view an approved external reference beside the candidate where applicable;
1342
- - inspect known observed regions;
1343
- - inspect known reference regions and their runtime bindings;
1344
- - see geometry;
1345
- - see relevant browser properties;
1346
- - inspect relationships;
1347
- - inspect before/after comparisons;
1348
- - inspect reference/candidate fidelity evidence;
1349
- - inspect contract results;
1350
- - understand warnings and failures.
1351
-
1352
- Selecting a structured runtime or reference region should identify the corresponding screenshot/reference area where practical.
1353
-
1354
- Likewise, selecting an image region should eventually support identifying the corresponding known runtime target or reference region when evidence is sufficient.
1355
-
1356
- The viewer must consume the reusable observation, reference, comparison, contract, correlation, and bounded-context engines/artifacts.
1357
-
1358
- It must not contain a second browser-observation implementation, a second reference model, a second binding engine, a second reference-evaluation implementation, a second contract engine, or a second bounded-context builder.
1359
-
1360
- ## Implemented capability (v0.9) — Human visual annotation
1361
-
1362
- Status: implemented and released as `0.9.0`. The intent below is unchanged and
1363
- remains the capability authority.
1364
-
1365
- A later phase should allow the user to communicate visual intent directly on top of either an observed frontend or an approved external reference.
1366
-
1367
- Useful annotation concepts may include:
1368
-
1369
- - freehand drawing;
1370
- - rectangle;
1371
- - arrow;
1372
- - line;
1373
- - textual note;
1374
- - preserve marker;
1375
- - resize marker;
1376
- - move marker;
1377
- - remove marker;
1378
- - inspect marker.
1379
-
1380
- Annotations must remain structured.
1381
-
1382
- Do not store annotation intent only as flattened image pixels.
1383
-
1384
- An annotation should preserve information such as:
1385
-
1386
- ```text
1387
- annotation source context: runtime observation or external reference
1388
- observation/screenshot identity or reference identity
1389
- annotation geometry
1390
- annotation type
1391
- textual instruction
1392
- associated runtime target/reference region where available
1393
- provenance and confirmation state
1394
- ```
1395
-
1396
- Example:
1397
-
1398
- ```text
1399
- annotation
1400
- → runtime target primary-navigation
1401
- → resize
1402
- → "make this visually narrower"
1403
- ```
1404
-
1405
- Another annotation may express:
1406
-
1407
- ```text
1408
- annotation
1409
- → reference region current-page-card
1410
- → "match this width and horizontal position"
1411
- ```
1412
-
1413
- Another annotation may express:
1414
-
1415
- ```text
1416
- annotation
1417
- → right-ad-rail
1418
- → preserve
1419
- ```
1420
-
1421
- The intended LLM-facing package may eventually combine:
1422
-
1423
- ```text
1424
- original screenshot
1425
- + approved reference image where applicable
1426
- + annotated screenshot/reference
1427
- + structured runtime observations
1428
- + structured reference evidence
1429
- + structured annotations
1430
- + current change scope
1431
- + previously approved contracts
1432
- ```
1433
-
1434
- This allows an LLM to reason simultaneously about:
1435
-
1436
- ```text
1437
- what exists
1438
- ```
1439
-
1440
- and:
1441
-
1442
- ```text
1443
- what the user wants changed or matched
1444
- ```
1445
-
1446
- Runtime-screenshot annotations and reference-image annotations remain different coordinate/identity domains. Ambiguous drawings must not silently become executable requirements.
1447
-
1448
- ## Relationship to `my-dev-kit`
1449
-
1450
- `my-dev-kit` and `my-frontend-observer` are sibling evidence producers.
1451
-
1452
- Conceptually:
1453
-
1454
- ```text
1455
- my-dev-kit
1456
- → what source exists?
1457
- → how is the repository structured?
1458
- → what symbols and dependencies matter?
1459
- → what source probably owns this behavior?
1460
- → what bounded source should the agent inspect?
1461
-
1462
- my-frontend-observer
1463
- → what did the browser actually render?
1464
- → where are the important regions?
1465
- → how large are they?
1466
- → what relationships exist?
1467
- → what is clipped or overflowing?
1468
- → what owns scrolling?
1469
- → what changed?
1470
- → what approved external reference should this candidate match?
1471
- → where does the candidate differ from that explicit reference intent?
1472
- ```
1473
-
1474
- Neither project should normally import or execute the other merely to perform its native responsibility.
1475
-
1476
- Their evidence may be correlated by an explicit consumer or integration contract.
1477
-
1478
- ## Relationship to `my-dev-kit-orchestrator`
1479
-
1480
- `my-dev-kit-orchestrator` owns workflow coordination rather than runtime observation or reference interpretation.
1481
-
1482
- The observer's v0.6/v0.7 public programmatic boundaries already expose bounded
1483
- runtime, correlation, fidelity, and correction-handoff evidence suitable for an
1484
- external orchestrator or coding-agent workflow. Orchestrator-side integration
1485
- remains a sibling-repository responsibility rather than code owned by this
1486
- repository.
1487
-
1488
- The orchestrator should not:
1489
-
1490
- - own browser automation;
1491
- - reproduce observer measurements;
1492
- - create its own external-reference schema;
1493
- - recompute reference/candidate fidelity;
1494
- - embed full raw observation/reference artifacts into prompts by default;
1495
- - redefine observer evidence semantics;
1496
- - become the canonical owner of observer artifacts.
1497
-
1498
- The observer exposes machine-consumable artifacts and a clean programmatic boundary so orchestrator integration does not require parsing human console output.
1499
-
1500
- ## Relationship to `my-dev-kit-lab`
1501
-
1502
- `my-dev-kit-lab` should evaluate observer compatibility and ecosystem behavior when coordinated validation requires it.
1503
-
1504
- Possible responsibilities include:
1505
-
1506
- - exact readers for supported observer artifact versions;
1507
- - pinned observer fixtures;
1508
- - browser/schema compatibility matrices;
1509
- - static/runtime correlation experiments;
1510
- - external-reference fixture and reader compatibility where required;
1511
- - reference/candidate evidence-quality evaluation where required;
1512
- - evidence-quality evaluation;
1513
- - controlled compatibility tests across ecosystem projects.
1514
-
1515
- The lab must not become the observer's production runtime or reference-evaluation engine.
1516
-
1517
- Normal frontend observation and normal reference-driven correction should not require the lab.
1518
-
1519
- ## Ecosystem integration principle
1520
-
1521
- Deep integration means:
1522
-
1523
- ```text
1524
- shared contracts
1525
- + explicit evidence boundaries
1526
- + compatible identities
1527
- + exact readers/adapters
1528
- + coordinated workflows
1529
- ```
1530
-
1531
- It does not mean:
1532
-
1533
- ```text
1534
- one package
1535
- one runtime
1536
- one schema for everything
1537
- or duplicated responsibilities
1538
- ```
1539
-
1540
- Do not introduce a shared cross-repository schema package merely for symmetry.
1541
-
1542
- A shared package should exist only if a future concrete integration demonstrates that it is necessary.
1543
-
1544
- ## Local-first requirement
1545
-
1546
- The tool should be local-first.
1547
-
1548
- The normal initial workflow should operate against applications running on:
1549
-
1550
- ```text
1551
- localhost
1552
- 127.0.0.1
1553
- local development hosts
1554
- ```
1555
-
1556
- Observation and reference-driven evaluation must not require uploading:
1557
-
1558
- - screenshots;
1559
- - external reference images;
1560
- - page contents;
1561
- - source code;
1562
- - observation artifacts;
1563
- - reference artifacts;
1564
- - visual annotations.
1565
-
1566
- No external artificial-intelligence API is required for the core observer.
1567
-
1568
- An LLM consuming generated evidence may operate separately from the observer.
1569
-
1570
- ## Browser and network safety
1571
-
1572
- Running a browser introduces a security and privacy boundary that must be defined explicitly.
1573
-
1574
- Before broad navigation support is implemented, the project must define behavior for matters such as:
1575
-
1576
- - allowed URL schemes;
1577
- - local versus remote targets;
1578
- - redirects;
1579
- - navigation timeouts;
1580
- - certificate failures;
1581
- - downloads;
1582
- - popups;
1583
- - browser permissions;
1584
- - network requests;
1585
- - unexpected navigation;
1586
- - credential-bearing pages;
1587
- - sensitive rendered data;
1588
- - secret-bearing URLs or output;
1589
- - cleanup of browser processes and temporary state.
1590
-
1591
- The first version should remain intentionally conservative and local.
1592
-
1593
- Safety behavior must be explicit rather than dependent on undocumented browser defaults.
1594
-
1595
- External-reference support has a separate local-file privacy boundary. v0.7
1596
- detects PNG/JPEG/WebP from header bytes, reads dimensions from bounded header
1597
- bytes without decoding pixels, enforces file-size/dimension bounds, keeps
1598
- operational paths out of semantic identity, and treats reference files as data,
1599
- not executable instructions.
1600
-
1601
- ## Target immutability
1602
-
1603
- Observation is non-destructive by default.
1604
-
1605
- The observer must not:
1606
-
1607
- - edit the target application's files;
1608
- - modify target source code;
1609
- - commit target changes;
1610
- - install dependencies into the target;
1611
- - alter target configuration;
1612
- - persist unintended application state;
1613
- - perform destructive interactions merely to collect layout evidence.
1614
-
1615
- The observed project is a target, not part of the observer repository.
1616
-
1617
- Interactions such as:
1618
-
1619
- - navigation;
1620
- - viewport resize;
1621
- - scrolling;
1622
- - explicitly approved safe controls;
1623
-
1624
- are acceptable when they are part of a defined observation scenario.
1625
-
1626
- A coding agent or another external tool performs source changes.
1627
-
1628
- Importing or evaluating an external reference does not authorize observer source edits or changes to the target application.
1629
-
1630
- ## Preferred platform
1631
-
1632
- Primary development platform:
1633
-
1634
- - desktop developer workstation;
1635
- - Windows first-class.
1636
-
1637
- The implementation must avoid unnecessary Windows-specific assumptions.
1638
-
1639
- Artifact paths, serialization, tests, and browser behavior should be designed so future ecosystem releases can satisfy the cross-platform validation expectations used by the broader `my-dev-kit` ecosystem.
1640
-
1641
- Cross-platform screenshot byte identity should not be assumed unless explicitly established by testing.
1642
-
1643
- Structured semantic evidence should remain the primary portable contract.
1644
-
1645
- The same caution applies to reference/candidate image similarity: font rasterization, graphics environment, antialiasing, and browser/OS differences must not be treated as exact semantic equality unless explicitly proven.
1646
-
1647
- ## Preferred implementation stack
1648
-
1649
- Preferred language:
1650
-
1651
- TypeScript.
1652
-
1653
- Preferred runtime:
1654
-
1655
- Node.js.
1656
-
1657
- Preferred browser automation:
1658
-
1659
- Playwright.
1660
-
1661
- Initial public interface:
1662
-
1663
- command-line interface (CLI).
1664
-
1665
- The first scaffold should favor a TypeScript/Node.js command-line project rather than a web-application-first architecture.
1666
-
1667
- The browser-observation engine must remain independent of command-line formatting so it can later support:
1668
-
1669
- - command-line use;
1670
- - programmatic use;
1671
- - graphical local viewing;
1672
- - automated regression workflows;
1673
- - ecosystem adapters.
1674
-
1675
- A later interactive viewer may use React or another suitable web user-interface stack.
1676
-
1677
- Do not put Playwright/browser-control logic directly inside React presentation components.
1678
-
1679
- Do not put canonical reference-evaluation logic directly inside viewer presentation components either.
1680
-
1681
- Do not promise a stable public programmatic application programming interface merely because internal modules are reusable.
1682
-
1683
- A public programmatic interface should become a compatibility commitment only when explicitly designed and tested.
1684
-
1685
- ## Architectural direction
1686
-
1687
- Use explicit ownership boundaries.
1688
-
1689
- The smallest expected conceptual separation is:
1690
-
1691
- ```text
1692
- command-line interface
1693
- ↓
1694
- observation application/engine
1695
- ↓
1696
- browser adapter
1697
- ↓
1698
- runtime evidence
1699
-
1700
- observation domain/schema
1701
- ↓
1702
- artifact writer
1703
-
1704
- deterministic fixture infrastructure
1705
- ↓
1706
- browser-level validation
1707
- ```
1708
-
1709
- The architecture has added, and later capabilities may continue to add:
1710
-
1711
- ```text
1712
- relationship engine
1713
- comparison engine
1714
- contract engine
1715
- bounded agent-context and correlation/export boundary
1716
- coding-agent review workflow
1717
- external visual-reference artifact/identity boundary
1718
- reference-to-runtime binding
1719
- reference/candidate structured evaluation
1720
- viewer
1721
- annotation system
1722
- ```
1723
-
1724
- These should extend the existing evidence model rather than creating parallel implementations.
1725
-
1726
- Avoid speculative abstraction.
1727
-
1728
- Do not create:
1729
-
1730
- - a generic plugin framework without multiple real implementations;
1731
- - a second observation engine for the viewer;
1732
- - a second comparison implementation for the user interface;
1733
- - a second contract engine for automated tests;
1734
- - a viewer-only reference model or reference-evaluation engine;
1735
- - annotation-only change semantics;
1736
- - a generic ecosystem evidence framework before concrete integration requires one.
1737
-
1738
- ## Initial product interface
1739
-
1740
- The initial public interface is CLI-first.
1741
-
1742
- The developer should be able to provide:
1743
-
1744
- ```text
1745
- target URL
1746
- viewport
1747
- observation targets
1748
- output location
1749
- ```
1750
-
1751
- and receive:
1752
-
1753
- ```text
1754
- screenshot
1755
- structured page observation
1756
- structured target observations
1757
- versioned observation artifact
1758
- concise execution/result summary
1759
- ```
1760
-
1761
- The CLI should be suitable for both human and machine invocation.
1762
-
1763
- Its architecture should leave room for:
1764
-
1765
- - machine-readable output;
1766
- - stable diagnostic codes;
1767
- - explicit exit behavior;
1768
- - separation between parseable output and human progress/diagnostics.
1769
-
1770
- The CLI must not own browser logic directly.
1771
-
1772
- The bounded agent context, static/runtime integration, text-driven coding-agent review, and non-graphical external-reference evidence foundation are on the core path after comparison/contracts. The graphical viewer and annotation system follow as human-interface enhancements.
1773
-
1774
- ## Evidence boundedness
1775
-
1776
- The observer must avoid collecting enormous amounts of runtime information merely because the browser exposes it.
1777
-
1778
- Initial observation should be explicitly scoped.
1779
-
1780
- Prefer:
1781
-
1782
- ```text
1783
- explicit observation targets
1784
- + required page facts
1785
- + required target facts
1786
- ```
1787
-
1788
- over:
1789
-
1790
- ```text
1791
- entire DOM
1792
- + every style property
1793
- + complete accessibility tree
1794
- ```
1795
-
1796
- Where evidence is bounded or truncated, the result should make the omission visible.
1797
-
1798
- A bounded collection should expose enough information to distinguish:
1799
-
1800
- ```text
1801
- nothing existed
1802
- ```
1803
-
1804
- from:
1805
-
1806
- ```text
1807
- evidence existed but was omitted because of a limit
1808
- ```
1809
-
1810
- Required evidence adequacy must not mean merely that some evidence was captured.
1811
-
1812
- If required configured evidence is missing, partial, or unavailable, the observer must say so.
1813
-
1814
- The same rule applies to references. Do not send every region, pixel delta, style sample, or image byte to a coding agent when only a bounded subset is relevant to the requested correction. Reference artifacts and bounded fidelity context expose omission/truncation where limits matter.
1815
-
1816
- ## Evidence provenance
1817
-
1818
- Runtime evidence should remain traceable to its source.
1819
-
1820
- Observation artifacts should record appropriate provenance such as:
1821
-
1822
- - observer package version;
1823
- - artifact schema version;
1824
- - browser engine;
1825
- - browser version;
1826
- - target URL;
1827
- - final URL;
1828
- - viewport;
1829
- - observation configuration;
1830
- - target identity and locator;
1831
- - observation method;
1832
- - artifact references;
1833
- - diagnostics;
1834
- - limits and omissions;
1835
- - capture identity;
1836
- - derivation method for derived facts.
1837
-
1838
- Reference evidence likewise preserves, as applicable:
1839
-
1840
- - exact reference identity/version;
1841
- - image reference and format/dimensions;
1842
- - region identity and coordinate semantics;
1843
- - authored requirements versus derived relationships;
1844
- - applicability state;
1845
- - approval/supersession state;
1846
- - binding evidence;
1847
- - tolerance/evaluation policy;
1848
- - candidate observation identity;
1849
- - diagnostics, limits, and omissions.
1850
-
1851
- Naturally unstable metadata such as capture time should not become the only logical identity of an observation or reference.
1852
-
1853
- ## Diagnostic behavior
1854
-
1855
- Observation failure and partial evidence must be explainable.
1856
-
1857
- The project should establish stable machine-readable diagnostics for cases such as:
1858
-
1859
- - invalid request;
1860
- - unsupported configuration;
1861
- - navigation failure;
1862
- - missing target;
1863
- - ambiguous target;
1864
- - hidden target;
1865
- - unavailable browser evidence;
1866
- - bounded/truncated evidence;
1867
- - artifact write failure;
1868
- - browser failure.
1869
-
1870
- Current reference support likewise makes malformed/unsupported reference data,
1871
- ambiguous/unavailable reference-to-runtime binding, incompatible
1872
- reference/candidate state, unavailable candidate evidence, and bounded fidelity
1873
- omissions explicit rather than fabricating normal values.
1874
-
1875
- Do not silently select an arbitrary target when selection is ambiguous.
1876
-
1877
- Do not represent unavailable evidence as a normal false or zero value.
1878
-
1879
- Warnings, partial observations, invalid requests, and fatal failures must remain distinguishable.
1880
-
1881
- ## Testing expectations
1882
-
1883
- Testing is a core requirement.
1884
-
1885
- The project should progressively include:
1886
-
1887
- ```text
1888
- unit tests
1889
- → schema/serialization tests
1890
- → browser adapter integration tests
1891
- → deterministic browser fixture tests
1892
- → comparison tests
1893
- → contract tests
1894
- → bounded agent-context and correlation tests
1895
- → ecosystem compatibility fixtures
1896
- → text-driven coding-agent workflow tests
1897
- → external-reference artifact/identity/binding tests
1898
- → reference applicability and structured fidelity tests
1899
- → reference-driven coding-agent correction tests
1900
- → viewer tests
1901
- → annotation tests for runtime and reference contexts
1902
- → full visual workflow tests
1903
- ```
1904
-
1905
- Important deterministic fixture scenarios should eventually include:
1906
-
1907
- - normal desktop layout;
1908
- - narrow navigation;
1909
- - clipped navigation contents;
1910
- - horizontal page overflow;
1911
- - nested scrolling;
1912
- - document scrolling;
1913
- - footer after workspace;
1914
- - overlapping regions;
1915
- - mobile layout;
1916
- - hidden elements;
1917
- - expected dependent resizing;
1918
- - protected-region regression;
1919
- - external reference whose candidate has measurable geometry/spacing mismatch;
1920
- - external reference with wrong theme/application-state candidate producing incompatibility;
1921
- - asset-sensitive reference region;
1922
- - reference-fidelity success coexisting with a protected-contract failure.
1923
-
1924
- The first version should use controlled local fixture pages rather than depending on public internet pages for canonical test evidence.
1925
-
1926
- Tests must distinguish:
1927
-
1928
- ```text
1929
- direct browser observation
1930
- derived interpretation
1931
- authored reference requirement
1932
- ```
1933
-
1934
- Screenshot evidence should not be treated as the only source of truth.
1935
-
1936
- Cross-platform tests should distinguish semantic/layout evidence from rendering differences that may legitimately vary by operating system, browser build, fonts, or graphics environment.
1937
-
1938
- ## Validation expectations
1939
-
1940
- The project should maintain a trustworthy validation chain appropriate to its current capabilities.
1941
-
1942
- At minimum, once established:
1943
-
1944
- ```text
1945
- typecheck
1946
- lint
1947
- unit/integration tests
1948
- browser fixture tests
1949
- production build when a graphical interface exists
1950
- documentation checks when implemented
1951
- ```
1952
-
1953
- Browser-related functionality must always have browser-level evidence.
1954
-
1955
- Passing static TypeScript validation alone is not sufficient for a browser-observation feature or a reference-driven workflow whose candidate side is browser-rendered.
1956
-
1957
- Later ecosystem releases should also satisfy the coordinated compatibility and cross-platform validation expectations of the `my-dev-kit` ecosystem.
1958
-
1959
- ## Performance expectations
1960
-
1961
- The tool is a developer utility.
1962
-
1963
- Correctness, boundedness, determinism, and inspectability are more important than extreme runtime optimization.
1964
-
1965
- However:
1966
-
1967
- - do not capture the entire Document Object Model when targeted evidence is sufficient;
1968
- - do not emit enormous computed-style dumps;
1969
- - do not take unnecessary screenshots;
1970
- - do not repeatedly decode/copy the same reference image into every downstream artifact;
1971
- - do not keep browser processes alive indefinitely;
1972
- - make observation and reference scope explicit;
1973
- - preserve evidence needed to explain conclusions;
1974
- - avoid duplicating unchanged evidence unnecessarily.
1975
-
1976
- ## Accessibility evidence
1977
-
1978
- Where the browser exposes it reliably, capture useful semantic/accessibility information such as:
1979
-
1980
- - role;
1981
- - accessible name;
1982
- - landmark identity;
1983
- - relevant state.
1984
-
1985
- This can help a human or LLM identify regions more reliably than position alone. Reference-to-runtime binding remains explicit and is never inferred merely because semantic evidence looks similar.
1986
-
1987
- The project is not initially intended to replace a dedicated accessibility-audit product.
1988
-
1989
- ## Inspectability
1990
-
1991
- Observation and regression results must be explainable.
1992
-
1993
- A useful result should identify:
1994
-
1995
- ```text
1996
- what was observed or explicitly referenced
1997
- where it was observed/referenced
1998
- what changed or still differs
1999
- before/reference value
2000
- after/candidate value
2001
- difference
2002
- expected condition
2003
- actual condition
2004
- contract, reference requirement, or relationship involved
2005
- supporting artifact
2006
- supporting screenshot/reference
2007
- ```
2008
-
2009
- Avoid unexplained scores.
2010
-
2011
- Avoid opaque artificial-intelligence classification in the core validation path.
2012
-
2013
- An LLM may reason over the evidence, but the evidence producer itself should remain inspectable.
2014
-
2015
- ## Determinism
2016
-
2017
- Given:
2018
-
2019
- - the same target build;
2020
- - the same browser version;
2021
- - the same viewport;
2022
- - the same observation configuration;
2023
- - the same deterministic fixture state;
2024
-
2025
- the structured observation should be stable enough for meaningful comparison.
2026
-
2027
- Given the same approved reference content, reference configuration, region definitions, applicability identity, and tolerance policy, the reference's logical identity and fidelity evaluation are deterministic according to the v0.7 contract.
2028
-
2029
- Fields that are naturally unstable must either:
2030
-
2031
- - be normalized;
2032
- - be excluded from logical comparison;
2033
- - or be explicitly identified as unstable metadata.
2034
-
2035
- Deterministic target ordering, reference-region ordering, diagnostic ordering, serialization, and artifact references should be preferred where practical.
2036
-
2037
- ## Non-goals for the initial project
2038
-
2039
- The initial project is not:
2040
-
2041
- - a replacement for browser developer tools;
2042
- - a replacement for Playwright;
2043
- - a replacement for `my-dev-kit`;
2044
- - a replacement for `my-dev-kit-orchestrator`;
2045
- - a replacement for `my-dev-kit-lab`;
2046
- - an autonomous frontend designer;
2047
- - an autonomous coding agent;
2048
- - a visual website builder;
2049
- - a hosted screenshot service;
2050
- - a cloud browser farm;
2051
- - a full accessibility scanner;
2052
- - a complete cross-browser testing service;
2053
- - a pixel-perfect visual-diff-only system;
2054
- - a Figma or Canva replacement;
2055
- - a screenshot-cloning SaaS;
2056
- - an autonomous raster-to-HTML/CSS generator;
2057
- - an automatic logo/vector reconstruction system;
2058
- - a general-purpose computer-vision framework;
2059
- - a source-code editor;
2060
- - a deployment system.
2061
-
2062
- The initial project does not need:
2063
-
2064
- - authentication;
2065
- - payments;
2066
- - advertising;
2067
- - multi-user collaboration;
2068
- - cloud persistence;
2069
- - remote browser infrastructure;
2070
- - external LLM APIs;
2071
- - production hosting;
2072
- - Firefox or WebKit support;
2073
- - source ownership;
2074
- - static repository indexing;
2075
- - orchestrator integration;
2076
- - lab integration;
2077
- - external visual-reference evaluation;
2078
- - visual annotation;
2079
- - comparison;
2080
- - regression contracts.
2081
-
2082
- Those capabilities may appear later according to Project Milestones and `ROADMAP.md`.
2083
-
2084
- ## Explicit product principles
2085
-
2086
- 1. Observe before inferring.
2087
- 2. Browser runtime is authoritative for rendered geometry.
2088
- 3. Source code and rendered output are different evidence domains.
2089
- 4. External desired-design references are a third evidence domain, distinct from runtime observations and source code.
2090
- 5. `my-dev-kit` owns static repository/source evidence; `my-frontend-observer` owns runtime browser evidence and its structured reference-evidence boundary.
2091
- 6. Stable runtime-region identity does not automatically imply known source ownership.
2092
- 7. Reference-region identity does not automatically imply runtime-target identity or source ownership.
2093
- 8. Observed dimensions are measurements, not automatically fixed design constants.
2094
- 9. A visible reference pixel is not automatically a hard requirement.
2095
- 10. Prefer relationship-based layout requirements when they better represent user intent.
2096
- 11. Distinguish direct browser facts, direct image measurements, authored requirements, and derived interpretations.
2097
- 12. Preserve raw evidence behind normalized and summarized evidence.
2098
- 13. Keep evidence bounded and make omissions explicit.
2099
- 14. Never represent unavailable evidence as if it were an observed false or zero.
2100
- 15. Never claim a visual requirement passed solely because a styling declaration looks correct.
2101
- 16. Never claim reference fidelity from pixel similarity alone when structured evidence is available or required.
2102
- 17. A requested change may legitimately cause dependent changes.
2103
- 18. Distinguish requested changes, expected dependent changes, protected properties, preserved invariants, and unexpected changes.
2104
- 19. A local requested change or reference match does not authorize unrelated rendered changes.
2105
- 20. Previously approved frontend invariants remain active unless the user explicitly supersedes them.
2106
- 21. Imported reference, approved reference, baseline approval, reference supersession, and baseline supersession are separate states/actions.
2107
- 22. Make regressions and fidelity failures explainable.
2108
- 23. Keep observation and reference evaluation non-destructive.
2109
- 24. Keep artifacts local-first, versioned, portable, and inspectable.
2110
- 25. Separate browser observation from static source analysis.
2111
- 26. Separate evidence production from workflow orchestration and downstream evaluation.
2112
- 27. Human visual intent must eventually be representable alongside machine measurements and external desired-design references.
2113
- 28. Deep ecosystem integration should use explicit contracts and adapters rather than duplicated responsibilities.
2114
- 29. Do not introduce speculative cross-project coupling before a real consumer requires it.
2115
- 30. The viewer and annotation layers consume canonical reference/comparison/contract evidence; they do not redefine it.
2116
-
2117
- ## Documentation and planning principles
2118
-
2119
- Documentation must distinguish current implemented behavior from future intended behavior.
2120
-
2121
- Current-state documentation should accurately record what exists.
2122
-
2123
- Forward-looking planning documents should preserve enough local design context for future LLM planning without requiring critical intent to be reconstructed from many unrelated bookkeeping documents.
2124
-
2125
- In particular:
2126
-
2127
- ```text
2128
- Project Description
2129
- → durable product intent
2130
- → responsibility boundaries
2131
- → long-term capability model
2132
-
2133
- Project Milestones
2134
- → ordered capability development
2135
- → major requirements
2136
- → acceptance expectations
2137
- → cross-milestone invariants
2138
-
2139
- ROADMAP.md
2140
- → high-level version specifications
2141
- → version goals
2142
- → required capabilities
2143
- → architectural constraints
2144
- → dependencies
2145
- → exclusions
2146
- → acceptance expectations
2147
- ```
2148
-
2149
- `ROADMAP.md` must not predefine implementation batches.
2150
-
2151
- When implementation of a roadmap version begins, the planner should:
2152
-
2153
- ```text
2154
- read the roadmap version
2155
- → inspect current repository state
2156
- → obtain required architecture/retrieval evidence
2157
- → design the implementation steps
2158
- → divide those steps into appropriate implementation batches
2159
- → execute and validate those batches
2160
- ```
2161
-
2162
- Forward-looking requirements may intentionally appear in more than one planning document when doing so prevents future planning context from becoming fragmented.
2163
-
2164
- ## Long-term product direction
2165
-
2166
- The long-term goal is to create a reliable communication and validation bridge between:
2167
-
2168
- ```text
2169
- human visual intent
2170
- approved external desired-design references where applicable
2171
- rendered frontend reality
2172
- static repository evidence
2173
- LLM reasoning
2174
- coding-agent implementation
2175
- ```
2176
-
2177
- The critical path has established:
2178
-
2179
- ```text
2180
- render and observe
2181
- → identify stable regions and runtime behavior
2182
- → compare
2183
- → enforce requested/dependent/protected/preserved scope
2184
- → combine bounded runtime and static evidence
2185
- → establish external-reference identity/regions/requirements/applicability/binding
2186
- → evaluate reference fidelity
2187
- → provide bounded context to an external coding agent
2188
- → rerender and reject fidelity or protected-contract regressions
2189
- ```
2190
-
2191
- Only after those core workflows work should the human visual branch add:
2192
-
2193
- ```text
2194
- viewer with reference/candidate inspection
2195
- → structured annotation on runtime or reference images
2196
- → full visual human–LLM workflow with both entry modes
2197
- ```
2198
-
2199
- The desired eventual visual cycle is:
2200
-
2201
- ```text
2202
- render/current candidate
2203
- + optional approved external reference
2204
- → observe
2205
- → identify stable runtime regions and reference regions
2206
- → measure geometry and behavior
2207
- → evaluate reference applicability/fidelity where applicable
2208
- → show human
2209
- → annotate/request change on runtime or reference
2210
- → define requested/dependent/protected/preserved scope
2211
- → combine bounded runtime/reference and static evidence
2212
- → provide context to LLM
2213
- → coding agent implements
2214
- → rerender
2215
- → compare before/after
2216
- → reevaluate reference/candidate fidelity
2217
- → rerun preserved contracts
2218
- → identify unexpected changes
2219
- → approve or correct
2220
- → establish new baseline and/or explicitly supersede reference according to policy
2221
- → repeat
2222
- ```
2223
-
2224
- The project succeeds when an LLM no longer needs to guess what a frontend looks like from source code alone, when a human can communicate visual intent or an approved desired design without translating every design idea into implementation terminology, and when a frontend change cannot be considered successful while silently breaking previously approved rendered behavior.
1
+ # my-frontend-observer
2
+
3
+ ## Project type
4
+
5
+ Greenfield developer tool and runtime-evidence producer within the `my-dev-kit` ecosystem.
6
+
7
+ ## Problem
8
+
9
+ Large language models (LLMs) and coding agents can inspect frontend source code, component trees, stylesheets, project architecture, and static dependencies, but they often cannot reliably understand what an application actually looks like or how it actually behaves after a browser renders it.
10
+
11
+ This creates a recurring frontend-development failure mode:
12
+
13
+ 1. A user describes a visual, layout, scrolling, responsiveness, or composition problem.
14
+ 2. The LLM interprets the request primarily through language and source code.
15
+ 3. Static repository evidence identifies a plausible source owner.
16
+ 4. A coding agent changes styling, layout, or component structure.
17
+ 5. The requested local symptom appears fixed.
18
+ 6. Another previously correct part of the rendered interface becomes visually or behaviorally broken.
19
+ 7. Source-level tests may still pass because the regression exists only in actual browser geometry, scrolling, overflow, clipping, spacing, responsiveness, or composition.
20
+ 8. The coding agent may incorrectly declare success because the requested source-level change was made without verifying the complete rendered result.
21
+
22
+ Examples include:
23
+
24
+ - shrinking a navigation column while leaving its contents too large for the new width;
25
+ - shrinking a navigation column without transferring the released space to the intended workspace;
26
+ - accidentally allowing an advertising rail to absorb released width;
27
+ - changing one grid track while unintentionally moving or resizing unrelated regions;
28
+ - fixing a nested scroll container in source while the rendered page still scrolls through the wrong container;
29
+ - causing labels to wrap, clip, overlap, or disappear;
30
+ - creating horizontal overflow at another viewport;
31
+ - moving, hiding, or resizing advertising, footer, header, or workspace regions unintentionally;
32
+ - satisfying one numerical styling requirement while degrading the composition as a whole;
33
+ - fixing one frontend problem while silently violating a previously approved frontend behavior.
34
+
35
+ Static repository understanding alone cannot reliably detect these failures because the authoritative evidence exists in the rendered browser.
36
+
37
+ There is also a communication problem.
38
+
39
+ A human often thinks about a frontend visually:
40
+
41
+ ```text
42
+ make this region narrower
43
+ move this boundary
44
+ give the released space to this region
45
+ preserve these regions
46
+ keep this scrolling behavior
47
+ do not change this layout relationship
48
+ ```
49
+
50
+ An LLM normally receives that intent as prose and must translate it into source changes without a reliable shared representation of the rendered interface.
51
+
52
+ A related failure occurs when the desired design already exists as an external visual reference. A user may have an approved PNG, WebP, screenshot, rendered mockup, or other image showing what the interface should look like, while the current application looks substantially different. Giving that image directly to a coding agent still leaves the agent to guess dimensions, spacing, region boundaries, relationships, state/theme applicability, style details, and which visual differences are actually requirements. Source/unit tests can pass while the result remains visibly far from the approved design.
53
+
54
+ The system therefore needs to represent both:
55
+
56
+ ```text
57
+ what the browser actually rendered
58
+ ```
59
+
60
+ and, when supplied:
61
+
62
+ ```text
63
+ what an approved external visual reference specifies
64
+ ```
65
+
66
+ without pretending that an external raster image is an earlier runtime observation or source-code artifact.
67
+
68
+ `my-frontend-observer` exists to provide that missing representation.
69
+
70
+ ## Product identity
71
+
72
+ `my-frontend-observer` is the runtime/browser evidence producer within the broader `my-dev-kit` ecosystem.
73
+
74
+ Its responsibility is:
75
+
76
+ ```text
77
+ running frontend
78
+ → real browser
79
+ → structured runtime evidence
80
+ ```
81
+
82
+ It also owns the structured evidence boundary for approved external visual references used to describe desired design intent. That reference evidence remains distinct from browser runtime evidence while being comparable to a rendered candidate through explicit bindings and evaluation.
83
+
84
+ It owns evidence about:
85
+
86
+ - what is actually rendered;
87
+ - where rendered regions are;
88
+ - how large they are;
89
+ - how they relate spatially;
90
+ - what is visible or clipped;
91
+ - what owns scrolling;
92
+ - whether overflow exists;
93
+ - what changed between observations;
94
+ - whether approved runtime relationships remain valid;
95
+ - what visual change the user intends;
96
+ - what an approved external visual reference specifies where one is supplied;
97
+ - which reference regions correspond to runtime targets where that binding is reliable;
98
+ - how a rendered candidate differs from explicit reference-design requirements.
99
+
100
+ It does not own static repository analysis.
101
+
102
+ The long-term ecosystem responsibility model is:
103
+
104
+ ```text
105
+ my-dev-kit
106
+ → static repository/source evidence producer
107
+ → files
108
+ → symbols
109
+ → dependencies
110
+ → architecture
111
+ → probable source ownership
112
+ → bounded source retrieval
113
+
114
+ my-frontend-observer
115
+ → rendered browser/runtime evidence producer
116
+ → screenshots
117
+ → rendered-region identity
118
+ → geometry
119
+ → layout relationships
120
+ → scrolling and overflow
121
+ → comparisons
122
+ → runtime contracts
123
+ → external visual-reference evidence
124
+ → reference/runtime binding
125
+ → visual intent
126
+
127
+ my-dev-kit-orchestrator
128
+ → coordinates development workflows
129
+ → consumes bounded evidence when appropriate
130
+ → prepares task context
131
+ → manages implementation/verification workflow
132
+
133
+ my-dev-kit-lab
134
+ → evaluates ecosystem behavior
135
+ → compatibility
136
+ → controlled fixtures
137
+ → experiments
138
+ → evidence quality
139
+ → cross-project validation
140
+ ```
141
+
142
+ These projects remain separately versioned and independently executable.
143
+
144
+ Deep ecosystem integration must not require collapsing their responsibilities into one package.
145
+
146
+ ## Product goal
147
+
148
+ Build `my-frontend-observer`, a local-first frontend observation, visual-communication, and runtime-regression tool that allows humans, LLMs, coding agents, and automated checks to reason from the frontend that the browser actually rendered and, when present, from an approved external visual reference describing the intended design.
149
+
150
+ The tool should convert browser state and explicitly supplied design-reference evidence into structured, inspectable evidence combining, as capabilities mature:
151
+
152
+ - screenshots;
153
+ - stable identities for meaningful rendered regions;
154
+ - relevant Document Object Model structure;
155
+ - rendered element geometry;
156
+ - computed browser layout properties;
157
+ - viewport information;
158
+ - scroll ownership and scroll state;
159
+ - visibility and overflow information;
160
+ - accessibility and semantic information;
161
+ - relationships between important rendered regions;
162
+ - before-and-after observations;
163
+ - persistent frontend invariants;
164
+ - requested-change intent;
165
+ - expected dependent changes;
166
+ - protected regions and properties;
167
+ - approved external visual references;
168
+ - stable identities for meaningful reference regions;
169
+ - reference geometry, relationships, applicability, tolerances, and provenance;
170
+ - explicit reference-region to runtime-target bindings;
171
+ - structured reference-design versus candidate evidence;
172
+ - structured visual annotations.
173
+
174
+ The primary goal is not to automatically redesign interfaces.
175
+
176
+ The primary goal is to give a human and an LLM a shared representation of:
177
+
178
+ > What is actually on the screen, where it is, how large it is, how it behaves, how its important regions relate to one another, what the user wants changed or wants matched from an approved reference, what must remain intact, and what actually changed after an implementation edit?
179
+
180
+ ## Three primary product jobs
181
+
182
+ ### 1. Human-to-LLM design and layout communication
183
+
184
+ The observer should help a user communicate visual and layout intent without requiring the LLM to infer everything from prose or source code.
185
+
186
+ The system should eventually allow communication through a combination of:
187
+
188
+ ```text
189
+ runtime screenshot or approved external reference
190
+ + stable named regions
191
+ + measured or authored geometry
192
+ + layout relationships
193
+ + runtime behavior where applicable
194
+ + visual annotation
195
+ + textual intent
196
+ ```
197
+
198
+ A user should be able to communicate ideas such as:
199
+
200
+ ```text
201
+ make primary navigation narrower
202
+ give the released horizontal space to the workspace
203
+ preserve both advertising rails
204
+ do not clip navigation labels
205
+ keep document-level scrolling
206
+ leave the footer relationship unchanged
207
+ ```
208
+
209
+ or:
210
+
211
+ ```text
212
+ make this popup match the approved One Dark reference
213
+ match this card's width and horizontal position
214
+ preserve these button proportions and spacing
215
+ use this approved artwork rather than redesigning it
216
+ this footer placement is informational, not an exact requirement
217
+ ```
218
+
219
+ without needing to express the implementation mechanism.
220
+
221
+ An external reference is desired-design evidence, not implementation. A raster image does not reveal original DOM structure, CSS, vector paths, component hierarchy, source ownership, hidden layout constraints, or design-token names unless those are separately supplied. The observer must preserve that distinction rather than fabricating hidden structure.
222
+
223
+ ### 2. Safe LLM-assisted frontend changes
224
+
225
+ The observer should make the complete rendered result part of the definition of implementation success.
226
+
227
+ A requested local change must not be considered successful merely because the requested element changed.
228
+
229
+ The tool should eventually distinguish:
230
+
231
+ ```text
232
+ requested change
233
+ expected dependent change
234
+ protected change
235
+ preserved invariant
236
+ unexpected change
237
+ ```
238
+
239
+ This creates an explicit allowed scope of frontend change.
240
+
241
+ Previously approved runtime behavior must remain valid unless the user explicitly supersedes it.
242
+
243
+ An external reference-fidelity success must obey the same rule. Matching a supplied design region does not authorize breaking a protected control, preserved relationship, baseline invariant, or unrelated application behavior.
244
+
245
+ ### 3. Runtime evidence for the my-dev-kit ecosystem
246
+
247
+ The observer should provide the runtime evidence domain that static repository analysis cannot provide.
248
+
249
+ The intended long-term combination is:
250
+
251
+ ```text
252
+ my-dev-kit static evidence
253
+ +
254
+ my-frontend-observer runtime evidence
255
+ +
256
+ my-frontend-observer reference evidence where applicable
257
+ ↓
258
+ bounded coordinated context
259
+ ↓
260
+ LLM / coding agent
261
+ ```
262
+
263
+ A runtime region may eventually be correlated with bounded source evidence without requiring the observer to become a source-analysis engine or requiring `my-dev-kit` to become a browser runner. Likewise, a reference region may bind to a runtime target without implying source ownership.
264
+
265
+ ## Intended users
266
+
267
+ Primary users:
268
+
269
+ - developers using LLMs or coding agents for frontend development;
270
+ - developers debugging visual and responsive regressions;
271
+ - developers who need to communicate visual layout intent to an LLM;
272
+ - developers who need to reproduce or validate an approved external visual design reference;
273
+ - developers who need browser-observed evidence before accepting frontend changes;
274
+ - maintainers who want machine-readable runtime frontend evidence;
275
+ - maintainers of the broader `my-dev-kit` ecosystem.
276
+
277
+ The initial user is a developer working locally with web applications, LLMs, and coding agents.
278
+
279
+ ## Initial workflow
280
+
281
+ The first useful workflow is intentionally small:
282
+
283
+ ```text
284
+ local web application
285
+ → target URL + viewport + explicit observation targets
286
+ → my-frontend-observer launches Chromium
287
+ → browser renders target
288
+ → observer captures screenshot
289
+ → observer captures structured page evidence
290
+ → observer captures structured target evidence
291
+ → observer writes a versioned local observation artifact
292
+ → command-line result reports completion, warnings, or failure
293
+ ```
294
+
295
+ The first version does not need comparison, change contracts, visual annotation, source ownership, external visual-reference evaluation, or direct integration with other ecosystem projects.
296
+
297
+ Its purpose is to establish a trustworthy runtime-evidence foundation.
298
+
299
+ ## Intended end-to-end workflow
300
+
301
+ The long-term workflow supports two human intent entry modes.
302
+
303
+ Actual-frontend-driven:
304
+
305
+ ```text
306
+ local web application
307
+ ↓
308
+ my-frontend-observer
309
+ ↓
310
+ baseline observation
311
+ + screenshot
312
+ + regions
313
+ + geometry
314
+ + relationships
315
+ + runtime behavior
316
+ ↓
317
+ human reviews frontend
318
+ ↓
319
+ human requests or visually annotates change
320
+ ```
321
+
322
+ Reference-driven:
323
+
324
+ ```text
325
+ approved external visual reference
326
+ + current browser-rendered candidate
327
+ ↓
328
+ reference identity, regions, relationships, applicability, and intent
329
+ ↓
330
+ explicit reference-region ↔ runtime-target binding
331
+ ↓
332
+ structured reference-vs-candidate evidence
333
+ ↓
334
+ human confirms which reference details are requirements
335
+ ```
336
+
337
+ Both then converge on:
338
+
339
+ ```text
340
+ requested / dependent / protected / preserved change scope
341
+ ↓
342
+ bounded runtime evidence
343
+ + relevant reference evidence where applicable
344
+ +
345
+ bounded static evidence from my-dev-kit where useful
346
+ ↓
347
+ my-dev-kit-orchestrator / developer / LLM
348
+ ↓
349
+ coding agent changes target source separately
350
+ ↓
351
+ my-frontend-observer captures new state
352
+ ↓
353
+ before/after comparison
354
+ + reference-vs-candidate evaluation where applicable
355
+ ↓
356
+ requested changes evaluated
357
+ + dependent changes evaluated
358
+ + protected properties evaluated
359
+ + existing regression contracts rerun
360
+ + reference requirements evaluated
361
+ ↓
362
+ PASS
363
+ or
364
+ actionable evidence identifying what broke or still differs
365
+ ↓
366
+ human approves new baseline/reference state
367
+ or requests another iteration
368
+ ```
369
+
370
+ The target application remains a separate project throughout this process.
371
+
372
+ ## Principal capability 1 — Browser observation
373
+
374
+ The tool must observe a locally running web application through a real browser.
375
+
376
+ Initial browser support should use Chromium through Playwright unless architecture work establishes a materially better supported mechanism.
377
+
378
+ The initial implementation should accept at minimum:
379
+
380
+ - target URL;
381
+ - viewport width;
382
+ - viewport height;
383
+ - explicitly configured observation targets;
384
+ - output location.
385
+
386
+ Later configuration may support:
387
+
388
+ - route collections;
389
+ - themes;
390
+ - reusable scenarios;
391
+ - browser-state setup;
392
+ - authentication setup;
393
+ - device profiles;
394
+ - interaction sequences.
395
+
396
+ Those later capabilities are not required for the first version.
397
+
398
+ Browser runtime behavior is authoritative for rendered geometry.
399
+
400
+ The observer must not infer final layout solely from source styles.
401
+
402
+ ## Principal capability 2 — Screenshot capture
403
+
404
+ For each observation, capture the rendered page as an image.
405
+
406
+ The screenshot is evidence associated with the same observation identity as the structured browser measurements.
407
+
408
+ Screenshots support:
409
+
410
+ - human review;
411
+ - multimodal LLM review;
412
+ - annotation;
413
+ - before/after inspection;
414
+ - reference/candidate inspection;
415
+ - regression evidence.
416
+
417
+ The system should eventually support:
418
+
419
+ - viewport screenshots;
420
+ - full-page screenshots where useful.
421
+
422
+ Exact initial screenshot behavior and capture-readiness semantics must be defined before implementation.
423
+
424
+ Pixel-perfect screenshot comparison must not become the only regression mechanism or the only reference-fidelity mechanism.
425
+
426
+ Structured browser evidence remains essential.
427
+
428
+ ## Principal capability 3 — Stable rendered-region identity
429
+
430
+ Meaningful rendered regions need stable logical identities so humans, LLMs, comparisons, annotations, regression contracts, and reference bindings can refer to the same conceptual runtime region over time.
431
+
432
+ Examples may include:
433
+
434
+ ```text
435
+ app-shell
436
+ header
437
+ primary-navigation
438
+ main-content
439
+ tool-workspace
440
+ left-ad-rail
441
+ right-ad-rail
442
+ footer-ad
443
+ footer
444
+ theme-control
445
+ ```
446
+
447
+ These names are examples only.
448
+
449
+ The observer must not assume that every application uses the same regions.
450
+
451
+ Region identity should support appropriate browser-observable mechanisms such as:
452
+
453
+ - semantic HTML elements;
454
+ - accessibility role;
455
+ - accessible name;
456
+ - stable `id`;
457
+ - stable `data-*` attribute;
458
+ - bounded CSS selector fallback;
459
+ - text-based selection only where appropriate.
460
+
461
+ A target may have a stable observer-level identity without having a known source-code component identity.
462
+
463
+ For example:
464
+
465
+ ```text
466
+ runtime target:
467
+ primary-navigation
468
+ ```
469
+
470
+ does not by itself prove:
471
+
472
+ ```text
473
+ source owner:
474
+ VerticalNav.tsx
475
+ ```
476
+
477
+ Source ownership belongs to the static-analysis integration boundary.
478
+
479
+ A reference region likewise has a separate reference identity. The released
480
+ v0.7 binding model keeps:
481
+
482
+ ```text
483
+ reference region:
484
+ primary-navigation-reference
485
+
486
+ runtime target:
487
+ primary-navigation
488
+ ```
489
+
490
+ as two explicit identity domains rather than collapsing them into one.
491
+
492
+ ## Principal capability 4 — Rendered layout map
493
+
494
+ Capture a structured representation of important rendered elements.
495
+
496
+ For an observed region, useful browser evidence includes:
497
+
498
+ ```text
499
+ identifier
500
+ selection method
501
+ semantic role
502
+ tag
503
+ accessible name where available
504
+ text summary where appropriate
505
+
506
+ x
507
+ y
508
+ width
509
+ height
510
+ right
511
+ bottom
512
+
513
+ visibility
514
+ display
515
+ position
516
+ overflow-x
517
+ overflow-y
518
+ z-index where relevant
519
+
520
+ scroll width
521
+ scroll height
522
+ client width
523
+ client height
524
+ scroll top
525
+ scroll left
526
+ ```
527
+
528
+ The observer should prefer browser-computed values over attempting to infer final geometry from source styling.
529
+
530
+ Observed dimensions are measurements, not automatically design constants.
531
+
532
+ For example:
533
+
534
+ ```text
535
+ primary-navigation.width = 176
536
+ ```
537
+
538
+ means:
539
+
540
+ ```text
541
+ the browser rendered the observed region at 176 pixels
542
+ ```
543
+
544
+ It does not automatically mean:
545
+
546
+ ```text
547
+ navigation must always be exactly 176 pixels wide
548
+ ```
549
+
550
+ Responsive layouts must remain possible.
551
+
552
+ The output must distinguish:
553
+
554
+ ```text
555
+ direct browser observation
556
+ computed browser property
557
+ derived relationship or interpretation
558
+ ```
559
+
560
+ ## Principal capability 5 — Page-level browser state
561
+
562
+ Capture page-level evidence such as:
563
+
564
+ ```text
565
+ URL
566
+ final URL after navigation
567
+ document title
568
+
569
+ viewport width
570
+ viewport height
571
+ device pixel ratio
572
+
573
+ document width
574
+ document height
575
+ document scroll width
576
+ document scroll height
577
+ document client width
578
+ document client height
579
+
580
+ window scroll X
581
+ window scroll Y
582
+
583
+ horizontal overflow state
584
+ vertical overflow state
585
+ ```
586
+
587
+ This should make questions such as these answerable from runtime evidence:
588
+
589
+ - Does the document own vertical scrolling?
590
+ - Is a child container actually scrolling instead?
591
+ - Is there horizontal document overflow?
592
+ - Is the footer below the initial viewport?
593
+ - Did the page become taller or wider after a change?
594
+ - Did viewport behavior change unexpectedly?
595
+
596
+ Reference-driven evaluation uses explicit caller-supplied state/applicability
597
+ identity so the observer does not compare the wrong theme, viewport,
598
+ authentication state, or application state as though it were the intended
599
+ reference state.
600
+
601
+ ## Principal capability 6 — Runtime scrolling, overflow, and visibility
602
+
603
+ Scrolling must be treated as runtime behavior rather than inferred solely from style declarations.
604
+
605
+ The observer should eventually be able to:
606
+
607
+ 1. capture initial scroll state;
608
+ 2. perform a controlled scroll action;
609
+ 3. capture resulting scroll state;
610
+ 4. identify which observed regions changed scroll position;
611
+ 5. expose evidence about which container appears to own scrolling;
612
+ 6. identify whether elements enter or leave the viewport;
613
+ 7. identify horizontal or vertical overflow.
614
+
615
+ Example direct observation:
616
+
617
+ ```text
618
+ before:
619
+ window.scrollY = 0
620
+ main.scrollTop = 0
621
+
622
+ after requested page scroll:
623
+ window.scrollY = 500
624
+ main.scrollTop = 0
625
+ ```
626
+
627
+ Possible derived interpretation:
628
+
629
+ ```text
630
+ document appears to own primary vertical scrolling
631
+ ```
632
+
633
+ The observer must not present the derived statement as if it were a direct browser measurement.
634
+
635
+ ## Principal capability 7 — Layout relationships and dependency relationships
636
+
637
+ Individual measurements are not enough.
638
+
639
+ Many design requirements concern relationships between regions.
640
+
641
+ The observer should support relationship-oriented evidence such as:
642
+
643
+ ```text
644
+ navigation is left of workspace
645
+ workspace is wider than navigation
646
+ navigation does not overlap workspace
647
+ workspace does not overlap right advertising rail
648
+ footer begins after main content
649
+ element is contained inside parent
650
+ navigation contents fit inside navigation
651
+ document width does not exceed viewport width
652
+ ```
653
+
654
+ The system should also leave room for an explicit layout relationship or dependency model.
655
+
656
+ Example:
657
+
658
+ ```text
659
+ Viewport
660
+ ↓
661
+ AppShell
662
+ ├── LeftAd
663
+ ├── Navigation
664
+ ├── Workspace
665
+ └── RightAd
666
+ ```
667
+
668
+ A requested change may imply legitimate dependent changes.
669
+
670
+ Example:
671
+
672
+ ```text
673
+ Navigation width decreases
674
+ ↓
675
+ Workspace width increases
676
+ Workspace x-position may move
677
+ ```
678
+
679
+ Other properties may need to remain preserved:
680
+
681
+ ```text
682
+ LeftAd width
683
+ RightAd width
684
+ Header height
685
+ Footer relationships
686
+ ```
687
+
688
+ The system must distinguish observed relationships from causal claims.
689
+
690
+ It should not automatically claim that one region caused another region to change merely because both changed.
691
+
692
+ Expected dependency semantics should come from an explicit contract, user intent, approved reference intent, or another supported source of evidence.
693
+
694
+ Reference regions reuse the canonical relationship vocabulary when the same
695
+ geometric relation applies, while preserving the fact that reference
696
+ relationships are derived from explicit reference-image geometry rather than
697
+ browser-observed DOM/runtime facts.
698
+
699
+ ## Principal capability 8 — Observation artifact
700
+
701
+ Each capture should produce one cohesive, observer-owned, versioned observation artifact or artifact directory.
702
+
703
+ The exact schema and filenames must be decided during architecture and schema design.
704
+
705
+ A conceptual structure may resemble:
706
+
707
+ ```text
708
+ observation/
709
+ manifest.json
710
+ page.json
711
+ elements.json
712
+ screenshot.png
713
+ ```
714
+
715
+ Possible future additions may include:
716
+
717
+ ```text
718
+ relationships.json
719
+ interactions.json
720
+ comparison.json
721
+ contracts.json
722
+ annotations.json
723
+ summary.txt
724
+ ```
725
+
726
+ These names are conceptual rather than fixed requirements.
727
+
728
+ The public artifact contract should establish from the beginning:
729
+
730
+ ```text
731
+ artifact kind
732
+ schema version
733
+ observation identity
734
+ producer version
735
+ browser identity
736
+ request/configuration identity
737
+ provenance
738
+ artifact references
739
+ completion state
740
+ diagnostics
741
+ limits
742
+ truncation/omission reporting
743
+ ```
744
+
745
+ Artifact paths should be relative and portable where possible.
746
+
747
+ Heavy evidence such as screenshots should be referenced rather than embedded into unrelated structured records.
748
+
749
+ Consumers must be able to distinguish a completed observation from a partial or failed capture.
750
+
751
+ The artifact must distinguish:
752
+
753
+ ```text
754
+ observed evidence
755
+ derived evidence
756
+ unavailable evidence
757
+ not-applicable evidence
758
+ partial evidence
759
+ ```
760
+
761
+ The artifact schema should evolve intentionally and additively where compatible.
762
+
763
+ Package version and observation schema version must remain separate concepts.
764
+
765
+ The v0.7 `ExternalReferenceArtifact` is a separate evidence family. It does
766
+ not masquerade as an observation merely to reuse an existing serializer.
767
+
768
+ ## Principal capability 9 — Before/after comparison
769
+
770
+ The tool should compare two observations representing comparable logical frontend states.
771
+
772
+ Useful differences include:
773
+
774
+ ```text
775
+ element moved
776
+ element resized
777
+ element disappeared
778
+ element appeared
779
+ visibility changed
780
+ element became clipped
781
+ horizontal overflow appeared
782
+ vertical overflow changed
783
+ document size changed
784
+ scroll-owner evidence changed
785
+ relative position changed
786
+ layout relationship changed
787
+ ```
788
+
789
+ Comparison should produce structured evidence such as:
790
+
791
+ ```text
792
+ target
793
+ property or relationship
794
+ before value
795
+ after value
796
+ difference
797
+ classification
798
+ supporting observation identities
799
+ ```
800
+
801
+ Example:
802
+
803
+ ```text
804
+ Target: primary-navigation
805
+ Property: width
806
+ Before: 176
807
+ After: 97
808
+ Difference: -79
809
+ ```
810
+
811
+ The comparison engine should preserve references to before/after screenshots and underlying observations.
812
+
813
+ The tool should not rely solely on screenshot pixel differences.
814
+
815
+ Before/after comparison is not part of the first observation version.
816
+
817
+ The initial observation identity and provenance model must nevertheless preserve enough information to support future comparability decisions.
818
+
819
+ Before/after comparison remains conceptually distinct from reference-design
820
+ versus candidate evaluation. An external desired-state image is not an earlier
821
+ runtime state.
822
+
823
+ ## Principal capability 10 — Explicit change scope
824
+
825
+ A central long-term concept is the ability to represent what a requested frontend change is allowed to affect.
826
+
827
+ A change should be expressible through categories such as:
828
+
829
+ ### Requested changes
830
+
831
+ Properties or relationships explicitly intended to change.
832
+
833
+ Example:
834
+
835
+ ```text
836
+ primary-navigation.width
837
+ → decrease significantly
838
+ ```
839
+
840
+ ### Expected dependent changes
841
+
842
+ Properties expected to change as a legitimate consequence.
843
+
844
+ Example:
845
+
846
+ ```text
847
+ tool-workspace.width
848
+ → increase using released horizontal space
849
+
850
+ tool-workspace.x
851
+ → may move left
852
+ ```
853
+
854
+ ### Protected properties or regions
855
+
856
+ Properties expected to remain unchanged.
857
+
858
+ Example:
859
+
860
+ ```text
861
+ left-ad-rail.width
862
+ right-ad-rail.width
863
+ header.height
864
+ ```
865
+
866
+ ### Preserved invariants and behaviors
867
+
868
+ Previously correct relationships or behaviors that must remain true.
869
+
870
+ Example:
871
+
872
+ ```text
873
+ navigation contents remain unclipped
874
+ navigation does not overlap workspace
875
+ workspace does not overlap advertising rails
876
+ document does not horizontally overflow
877
+ document continues to own primary page scrolling
878
+ mobile layout remains usable
879
+ ```
880
+
881
+ Together, these categories define the allowed scope of rendered change.
882
+
883
+ This concept may eventually be represented by an explicit `ChangeContract` or equivalent schema.
884
+
885
+ The conceptual name does not require that exact implementation type.
886
+
887
+ Reference-derived executable intent maps into this same change-scope model. A
888
+ visible detail in a reference may remain informational or unassessed until
889
+ explicitly promoted into requested, expected-dependent, protected, or
890
+ preserved intent. There is no separate reference-only change taxonomy.
891
+
892
+ ## Principal capability 11 — Frontend regression and change contracts
893
+
894
+ The project should support persistent executable runtime invariants.
895
+
896
+ Examples include:
897
+
898
+ ```text
899
+ element is visible
900
+ element is not clipped
901
+ element width is within a bound
902
+ element A does not overlap element B
903
+ element A is wider than element B
904
+ element A follows element B vertically
905
+ document width does not exceed viewport width
906
+ window owns requested page scrolling
907
+ specified element does not own primary page scrolling
908
+ element begins below initial viewport
909
+ ```
910
+
911
+ Relationship-oriented contracts should be preferred when they represent user intent more accurately than fixed pixels.
912
+
913
+ For example:
914
+
915
+ Prefer:
916
+
917
+ ```text
918
+ workspace width increases when navigation width decreases
919
+ ```
920
+
921
+ when that is the actual design requirement.
922
+
923
+ Use:
924
+
925
+ ```text
926
+ navigation.width = 97
927
+ ```
928
+
929
+ only when the user truly requires that exact value.
930
+
931
+ The system should support two related forms of contract:
932
+
933
+ ```text
934
+ persistent baseline contracts
935
+ ```
936
+
937
+ and:
938
+
939
+ ```text
940
+ per-change contracts
941
+ ```
942
+
943
+ Persistent baseline contracts preserve approved frontend behavior across future changes.
944
+
945
+ Per-change contracts describe:
946
+
947
+ ```text
948
+ requested changes
949
+ expected dependent changes
950
+ protected properties
951
+ preserved invariants
952
+ ```
953
+
954
+ Example evaluation:
955
+
956
+ ```text
957
+ REQUESTED CHANGE
958
+ Navigation.width
959
+ 176 → 97
960
+ PASS
961
+
962
+ EXPECTED DEPENDENT CHANGE
963
+ Workspace.width
964
+ 960 → 1039
965
+ PASS
966
+
967
+ PROTECTED PROPERTY
968
+ RightAd.width
969
+ 112 → 154
970
+ FAIL
971
+
972
+ PRESERVED INVARIANT
973
+ Navigation content became clipped
974
+ FAIL
975
+
976
+ OVERALL
977
+ FAIL
978
+ ```
979
+
980
+ A frontend change must not be declared successful merely because its requested local mutation succeeded.
981
+
982
+ Reference fidelity supplements these contracts. It does not replace or weaken
983
+ them, and a fidelity pass cannot override a protected or preserved contract
984
+ failure.
985
+
986
+ ## Principal capability 12 — Bounded agent context and static/runtime integration
987
+
988
+ Structured output must support both programmatic use and LLM consumption.
989
+
990
+ The tool should eventually produce a bounded runtime-evidence package containing, as applicable:
991
+
992
+ - target page identity;
993
+ - viewport;
994
+ - major observed regions;
995
+ - region geometry;
996
+ - semantic identities;
997
+ - layout relationships;
998
+ - dependency/change-scope information;
999
+ - overflow state;
1000
+ - scroll evidence;
1001
+ - comparison results;
1002
+ - contract results;
1003
+ - important warnings;
1004
+ - references to underlying raw evidence.
1005
+
1006
+ Preserve the evidence hierarchy:
1007
+
1008
+ ```text
1009
+ raw browser evidence
1010
+ ↓
1011
+ normalized structured evidence
1012
+ ↓
1013
+ derived relationships
1014
+ ↓
1015
+ bounded summary/context
1016
+ ↓
1017
+ LLM reasoning
1018
+ ```
1019
+
1020
+ The bounded context must not require an LLM to consume:
1021
+
1022
+ - an entire raw Document Object Model dump;
1023
+ - every computed style property;
1024
+ - enormous accessibility trees;
1025
+ - repeated unchanged measurements;
1026
+ - every screenshot produced during a workflow.
1027
+
1028
+ The summary must remain traceable to the evidence supporting it.
1029
+
1030
+ The initial command-line version may return a concise execution summary.
1031
+
1032
+ That operational summary must not be confused with the richer agent-oriented context package described here.
1033
+
1034
+ The shortest path to practical coding-agent use combines this bounded runtime
1035
+ projection with relevant bounded static/source evidence from `my-dev-kit`.
1036
+ The evidence domains remain separate and traceable:
1037
+
1038
+ ```text
1039
+ observer runtime evidence
1040
+ +
1041
+ my-dev-kit static evidence
1042
+ ↓
1043
+ bounded agent context
1044
+ ↓
1045
+ external coding agent
1046
+ ```
1047
+
1048
+ When an external reference is active, the released v0.7 extension adds only
1049
+ task-relevant reference identity, selected design requirements, measurable
1050
+ candidate mismatches, bound runtime targets, protected/preserved context, and
1051
+ references to heavy image assets. It does not place the full reference artifact
1052
+ or every image difference into the agent packet by default.
1053
+
1054
+ Runtime/static correlation must be explicit and may be ambiguous. A stable
1055
+ runtime target identity must never silently become a source-ownership claim.
1056
+ The observer owns runtime projection and its correlation/export boundary;
1057
+ `my-dev-kit` owns static indexing and retrieval; the orchestrator coordinates
1058
+ bounded consumption; the lab owns exact compatibility evaluation.
1059
+
1060
+ This integrated, text/config-driven path supports an end-to-end coding-agent
1061
+ change review before the viewer or visual annotation becomes a prerequisite.
1062
+ The observer does not edit source: an external coding agent makes the change,
1063
+ after which the observer rerenders, compares, and evaluates preserved contracts
1064
+ and, where applicable, reference fidelity.
1065
+
1066
+ ## Implemented capability (v0.7) — External visual reference and reference-driven design evidence
1067
+
1068
+ v0.7 supports approved external visual references as a structured desired-design evidence domain.
1069
+
1070
+ The public `import-reference` command accepts PNG, JPEG, and WebP images. Format
1071
+ and dimensions are detected from bounded header bytes rather than trusted from a
1072
+ filename extension, and the implementation enforces bounded file-size and image-
1073
+ dimension limits. No OCR, image segmentation, computer-vision target discovery,
1074
+ or raster-to-code reconstruction is part of this capability.
1075
+
1076
+ A raw reference image is evidence, not implementation. It does not reveal hidden DOM structure, source ownership, original CSS, component hierarchy, design tokens, original vector paths, or inaccessible font metadata.
1077
+
1078
+ The `ExternalReferenceArtifact` and its derived evaluation path preserve enough
1079
+ structured information to answer:
1080
+
1081
+ ```text
1082
+ which exact reference image/version was used?
1083
+ which regions were defined?
1084
+ which requirements were explicitly authored?
1085
+ which relationships were derived?
1086
+ which candidate observation was evaluated?
1087
+ which viewport/theme/application state applies?
1088
+ which tolerance/evaluation policy was used?
1089
+ which approval or supersession decision applies?
1090
+ ```
1091
+
1092
+ ### Reference-region model
1093
+
1094
+ v0.7 reference regions are explicit, bounded semantic rectangles authored by a
1095
+ user or configuration. Each region has a stable `id` and a canonical
1096
+ `{x, y, width, height}` rectangle in reference-image pixels with origin at the
1097
+ image's top-left corner. `right`, `bottom`, `centerX`, and `centerY` are derived
1098
+ on demand from that canonical rectangle and are not redundantly persisted.
1099
+ There is no automatic segmentation and no normalized-coordinate region model in
1100
+ v0.7.
1101
+
1102
+ Reference regions reuse the same geometry-only relationship predicates used by
1103
+ runtime layout relationships where the concept is genuinely shared, including
1104
+ horizontal/vertical order, overlap, relative width, geometric fit, and vertical
1105
+ sequencing. Reference relationships are derived on demand and do not become
1106
+ requirements automatically.
1107
+
1108
+ ### Reference-design intent and tolerances
1109
+
1110
+ Not every visible pixel is a requirement.
1111
+
1112
+ Reference evidence distinguishes:
1113
+
1114
+ ```text
1115
+ visible/derived evidence
1116
+ explicit authored requirement
1117
+ informational/unassessed detail
1118
+ ```
1119
+
1120
+ Executable reference requirements reuse the canonical v0.5 authored categories:
1121
+
1122
+ ```text
1123
+ requested
1124
+ expected-dependent
1125
+ protected
1126
+ preserved
1127
+ ```
1128
+
1129
+ `unexpected` remains derived-only.
1130
+
1131
+ v0.7 supports selected requirement subjects over region properties,
1132
+ region-to-region relationships, and bounded two-region measurements. Numeric
1133
+ reference tolerances are explicitly reference-owned and use:
1134
+
1135
+ - `exact`;
1136
+ - `absolute-reference-px`;
1137
+ - `percent`.
1138
+
1139
+ Reference-image coordinates and tolerances are not silently treated as CSS
1140
+ pixels. Fidelity establishes an explicit full-frame reference-image-pixel to
1141
+ CSS-pixel scale from the reference image dimensions and declared applicable
1142
+ runtime viewport, with an independent aspect-ratio-coherence gate.
1143
+
1144
+ Selected color/style evidence, asset-similarity evidence, or image-region
1145
+ similarity are not v0.7 success mechanisms. They may be added later only as
1146
+ bounded supplemental evidence and must not replace structured geometry,
1147
+ relationships, applicability, or canonical contract evaluation.
1148
+
1149
+ ### Reference applicability and comparability
1150
+
1151
+ Reference and candidate must represent compatible intended states before ordinary fidelity differences are evaluated.
1152
+
1153
+ v0.7 supports explicit caller-supplied applicability dimensions for:
1154
+
1155
+ ```text
1156
+ viewport
1157
+ theme
1158
+ application state
1159
+ authenticated state
1160
+ ```
1161
+
1162
+ The corresponding candidate state is likewise caller/configuration supplied on
1163
+ observation. It is not inferred from screenshot pixels, DOM, CSS, URL, or source
1164
+ code.
1165
+
1166
+ For example:
1167
+
1168
+ ```text
1169
+ reference: One Dark / active crawl
1170
+ candidate: One Light / idle
1171
+ ```
1172
+
1173
+ produces an explicit incompatible/incomparable result rather than a meaningless visual-difference list when those dimensions are declared and conflict.
1174
+
1175
+ The compatibility implementation reuses the existing v0.4 comparability result
1176
+ and per-dimension comparison conventions rather than creating an unrelated
1177
+ reference-only state system.
1178
+
1179
+ ### Reference-to-runtime binding
1180
+
1181
+ The system uses an explicit association between:
1182
+
1183
+ ```text
1184
+ reference region
1185
+ ```
1186
+
1187
+ and:
1188
+
1189
+ ```text
1190
+ runtime target
1191
+ ```
1192
+
1193
+ Binding declarations are caller/configuration supplied and are never inferred
1194
+ from geometry, matching names, or source code. Binding results use the closed
1195
+ states:
1196
+
1197
+ ```text
1198
+ bound
1199
+ ambiguous
1200
+ unavailable
1201
+ ```
1202
+
1203
+ Reference identity, runtime identity, and source identity remain separate domains.
1204
+
1205
+ ### Structured reference-vs-candidate evaluation
1206
+
1207
+ v0.7 combines selected reference requirements with browser-authoritative
1208
+ candidate evidence and produces bounded, actionable structured fidelity results.
1209
+ Evaluation proceeds through reference structural validation, reference-evidence
1210
+ adequacy, reference/candidate compatibility, explicit binding, then each
1211
+ selected requirement. An inadequate reference or incompatible candidate is
1212
+ `not-evaluated`; it is not fabricated into an ordinary visual failure.
1213
+
1214
+ Example evidence remains of the form:
1215
+
1216
+ ```text
1217
+ Target: current-page-card
1218
+ Reference x: 28
1219
+ Candidate x: 18
1220
+ Delta: -10
1221
+
1222
+ Reference width: 424
1223
+ Candidate width: 446
1224
+ Delta: +22
1225
+
1226
+ Expected separation below header: 24px within tolerance
1227
+ Candidate separation: 38px
1228
+ Result: fidelity requirement failed
1229
+ ```
1230
+
1231
+ Pixel or image similarity is not the success mechanism in v0.7.
1232
+
1233
+ ### Reference lifecycle and approval
1234
+
1235
+ A random supplied image never silently becomes a project baseline or active design authority.
1236
+
1237
+ The v0.7 persisted lifecycle has exactly two explicit states:
1238
+
1239
+ ```text
1240
+ imported
1241
+ → approved
1242
+ ```
1243
+
1244
+ `import-reference` creates a new imported artifact. `approve-reference` is the
1245
+ only explicit approval act and creates a new approved artifact instance while
1246
+ preserving the imported artifact unchanged. Supersession is represented by a
1247
+ forward pointer on the newer artifact and never rewrites the superseded
1248
+ artifact. There is no automatic measured, annotated, active, or auto-approved
1249
+ lifecycle state in v0.7.
1250
+
1251
+ Reference approval is separate from baseline approval. Reference supersession is separate from baseline supersession. A fidelity `PASS` does not approve either one automatically.
1252
+
1253
+ ### Multiple references
1254
+
1255
+ The model permits separately identified references for explicit states such as:
1256
+
1257
+ ```text
1258
+ dark theme / idle
1259
+ dark theme / active
1260
+ dark theme / error
1261
+ light theme / idle
1262
+ light theme / active
1263
+ light theme / error
1264
+ desktop
1265
+ mobile
1266
+ ```
1267
+
1268
+ The correct reference must be selected by explicit identity/applicability rules rather than by accidental filename matching.
1269
+
1270
+ ### Asset fidelity
1271
+
1272
+ A reference region may represent artwork or another asset-sensitive area.
1273
+
1274
+ The observer can currently preserve the reference image and structured region
1275
+ geometry but does not claim to recover vector paths or hidden source data from a
1276
+ raster reference. Raster-to-vector reconstruction and image-to-code generation
1277
+ remain external implementation concerns. Bounded asset/image-similarity evidence
1278
+ would be a later extension, not a current v0.7 contract.
1279
+
1280
+ ### Coding-agent correction packet
1281
+
1282
+ The v0.7 bounded agent context reports measurable reference/candidate mismatches
1283
+ rather than asking the coding agent to reinterpret the entire image each
1284
+ iteration. It carries relevant failed requirements, bound runtime targets,
1285
+ active protected/preserved context, provenance, adequacy/omission/truncation,
1286
+ and bounded static/source correlation when supplied by the caller.
1287
+
1288
+ The implemented correction loop is:
1289
+
1290
+ ```text
1291
+ approved external reference
1292
+ → structured reference evidence
1293
+ → current candidate observation
1294
+ → structured fidelity mismatch
1295
+ → bounded runtime/static context
1296
+ → external coding agent correction
1297
+ → real Chromium rerender
1298
+ → reevaluate reference fidelity
1299
+ + canonical before/after comparison
1300
+ + canonical baseline/per-change contract evaluation
1301
+ → PASS or actionable failure
1302
+ ```
1303
+
1304
+ Matching the reference is necessary but never sufficient: an active protected or
1305
+ preserved contract regression still makes the overall correction review fail.
1306
+ The observer never edits target source; the implementation actor remains external.
1307
+
1308
+ The viewer and annotation systems later consume this reference model. They must not create another one.
1309
+
1310
+ ## Implemented foundation (v0.6) — Static/runtime source association
1311
+
1312
+ The observer provides an explicit programmatic runtime/static correlation
1313
+ boundary for associating stable runtime targets with caller-supplied bounded
1314
+ static candidates where reliable.
1315
+
1316
+ The chain is:
1317
+
1318
+ ```text
1319
+ rendered region
1320
+ → runtime target identity
1321
+ → correlation evidence
1322
+ → my-dev-kit static identity / bounded evidence
1323
+ → relevant source retrieval
1324
+ ```
1325
+
1326
+ Correlation results preserve `correlated`, `ambiguous`, or `unavailable`
1327
+ outcomes and competing candidates. The observer does not implement a competing
1328
+ repository-analysis system and does not silently turn a runtime target into a
1329
+ source owner. `my-dev-kit` remains the owner of repository crawling, parsing,
1330
+ indexing, source graphs, architecture, and bounded retrieval. The observer
1331
+ package has no runtime dependency on `@dailephd/my-dev-kit`; static candidate
1332
+ evidence is supplied through the explicit boundary.
1333
+
1334
+ ## Principal capability 13 — Human visual review
1335
+
1336
+ After the text/config-driven coding-agent workflow and non-graphical external-reference evidence foundation are proven, the project should provide a human-readable graphical way to inspect the same canonical evidence.
1337
+
1338
+ A later local interface should allow the developer to:
1339
+
1340
+ - view the captured screenshot;
1341
+ - view an approved external reference beside the candidate where applicable;
1342
+ - inspect known observed regions;
1343
+ - inspect known reference regions and their runtime bindings;
1344
+ - see geometry;
1345
+ - see relevant browser properties;
1346
+ - inspect relationships;
1347
+ - inspect before/after comparisons;
1348
+ - inspect reference/candidate fidelity evidence;
1349
+ - inspect contract results;
1350
+ - understand warnings and failures.
1351
+
1352
+ Selecting a structured runtime or reference region should identify the corresponding screenshot/reference area where practical.
1353
+
1354
+ Likewise, selecting an image region should eventually support identifying the corresponding known runtime target or reference region when evidence is sufficient.
1355
+
1356
+ The viewer must consume the reusable observation, reference, comparison, contract, correlation, and bounded-context engines/artifacts.
1357
+
1358
+ It must not contain a second browser-observation implementation, a second reference model, a second binding engine, a second reference-evaluation implementation, a second contract engine, or a second bounded-context builder.
1359
+
1360
+ ## Implemented capability (v0.9) — Human visual annotation
1361
+
1362
+ Status: implemented and released as `0.9.0`. The intent below is unchanged and
1363
+ remains the capability authority.
1364
+
1365
+ A later phase should allow the user to communicate visual intent directly on top of either an observed frontend or an approved external reference.
1366
+
1367
+ Useful annotation concepts may include:
1368
+
1369
+ - freehand drawing;
1370
+ - rectangle;
1371
+ - arrow;
1372
+ - line;
1373
+ - textual note;
1374
+ - preserve marker;
1375
+ - resize marker;
1376
+ - move marker;
1377
+ - remove marker;
1378
+ - inspect marker.
1379
+
1380
+ Annotations must remain structured.
1381
+
1382
+ Do not store annotation intent only as flattened image pixels.
1383
+
1384
+ An annotation should preserve information such as:
1385
+
1386
+ ```text
1387
+ annotation source context: runtime observation or external reference
1388
+ observation/screenshot identity or reference identity
1389
+ annotation geometry
1390
+ annotation type
1391
+ textual instruction
1392
+ associated runtime target/reference region where available
1393
+ provenance and confirmation state
1394
+ ```
1395
+
1396
+ Example:
1397
+
1398
+ ```text
1399
+ annotation
1400
+ → runtime target primary-navigation
1401
+ → resize
1402
+ → "make this visually narrower"
1403
+ ```
1404
+
1405
+ Another annotation may express:
1406
+
1407
+ ```text
1408
+ annotation
1409
+ → reference region current-page-card
1410
+ → "match this width and horizontal position"
1411
+ ```
1412
+
1413
+ Another annotation may express:
1414
+
1415
+ ```text
1416
+ annotation
1417
+ → right-ad-rail
1418
+ → preserve
1419
+ ```
1420
+
1421
+ The intended LLM-facing package may eventually combine:
1422
+
1423
+ ```text
1424
+ original screenshot
1425
+ + approved reference image where applicable
1426
+ + annotated screenshot/reference
1427
+ + structured runtime observations
1428
+ + structured reference evidence
1429
+ + structured annotations
1430
+ + current change scope
1431
+ + previously approved contracts
1432
+ ```
1433
+
1434
+ This allows an LLM to reason simultaneously about:
1435
+
1436
+ ```text
1437
+ what exists
1438
+ ```
1439
+
1440
+ and:
1441
+
1442
+ ```text
1443
+ what the user wants changed or matched
1444
+ ```
1445
+
1446
+ Runtime-screenshot annotations and reference-image annotations remain different coordinate/identity domains. Ambiguous drawings must not silently become executable requirements.
1447
+
1448
+ ## Relationship to `my-dev-kit`
1449
+
1450
+ `my-dev-kit` and `my-frontend-observer` are sibling evidence producers.
1451
+
1452
+ Conceptually:
1453
+
1454
+ ```text
1455
+ my-dev-kit
1456
+ → what source exists?
1457
+ → how is the repository structured?
1458
+ → what symbols and dependencies matter?
1459
+ → what source probably owns this behavior?
1460
+ → what bounded source should the agent inspect?
1461
+
1462
+ my-frontend-observer
1463
+ → what did the browser actually render?
1464
+ → where are the important regions?
1465
+ → how large are they?
1466
+ → what relationships exist?
1467
+ → what is clipped or overflowing?
1468
+ → what owns scrolling?
1469
+ → what changed?
1470
+ → what approved external reference should this candidate match?
1471
+ → where does the candidate differ from that explicit reference intent?
1472
+ ```
1473
+
1474
+ Neither project should normally import or execute the other merely to perform its native responsibility.
1475
+
1476
+ Their evidence may be correlated by an explicit consumer or integration contract.
1477
+
1478
+ ## Relationship to `my-dev-kit-orchestrator`
1479
+
1480
+ `my-dev-kit-orchestrator` owns workflow coordination rather than runtime observation or reference interpretation.
1481
+
1482
+ The observer's v0.6/v0.7 public programmatic boundaries already expose bounded
1483
+ runtime, correlation, fidelity, and correction-handoff evidence suitable for an
1484
+ external orchestrator or coding-agent workflow. Orchestrator-side integration
1485
+ remains a sibling-repository responsibility rather than code owned by this
1486
+ repository.
1487
+
1488
+ The orchestrator should not:
1489
+
1490
+ - own browser automation;
1491
+ - reproduce observer measurements;
1492
+ - create its own external-reference schema;
1493
+ - recompute reference/candidate fidelity;
1494
+ - embed full raw observation/reference artifacts into prompts by default;
1495
+ - redefine observer evidence semantics;
1496
+ - become the canonical owner of observer artifacts.
1497
+
1498
+ The observer exposes machine-consumable artifacts and a clean programmatic boundary so orchestrator integration does not require parsing human console output.
1499
+
1500
+ ## Relationship to `my-dev-kit-lab`
1501
+
1502
+ `my-dev-kit-lab` should evaluate observer compatibility and ecosystem behavior when coordinated validation requires it.
1503
+
1504
+ Possible responsibilities include:
1505
+
1506
+ - exact readers for supported observer artifact versions;
1507
+ - pinned observer fixtures;
1508
+ - browser/schema compatibility matrices;
1509
+ - static/runtime correlation experiments;
1510
+ - external-reference fixture and reader compatibility where required;
1511
+ - reference/candidate evidence-quality evaluation where required;
1512
+ - evidence-quality evaluation;
1513
+ - controlled compatibility tests across ecosystem projects.
1514
+
1515
+ The lab must not become the observer's production runtime or reference-evaluation engine.
1516
+
1517
+ Normal frontend observation and normal reference-driven correction should not require the lab.
1518
+
1519
+ ## Ecosystem integration principle
1520
+
1521
+ Deep integration means:
1522
+
1523
+ ```text
1524
+ shared contracts
1525
+ + explicit evidence boundaries
1526
+ + compatible identities
1527
+ + exact readers/adapters
1528
+ + coordinated workflows
1529
+ ```
1530
+
1531
+ It does not mean:
1532
+
1533
+ ```text
1534
+ one package
1535
+ one runtime
1536
+ one schema for everything
1537
+ or duplicated responsibilities
1538
+ ```
1539
+
1540
+ Do not introduce a shared cross-repository schema package merely for symmetry.
1541
+
1542
+ A shared package should exist only if a future concrete integration demonstrates that it is necessary.
1543
+
1544
+ ## Local-first requirement
1545
+
1546
+ The tool should be local-first.
1547
+
1548
+ The normal initial workflow should operate against applications running on:
1549
+
1550
+ ```text
1551
+ localhost
1552
+ 127.0.0.1
1553
+ local development hosts
1554
+ ```
1555
+
1556
+ Observation and reference-driven evaluation must not require uploading:
1557
+
1558
+ - screenshots;
1559
+ - external reference images;
1560
+ - page contents;
1561
+ - source code;
1562
+ - observation artifacts;
1563
+ - reference artifacts;
1564
+ - visual annotations.
1565
+
1566
+ No external artificial-intelligence API is required for the core observer.
1567
+
1568
+ An LLM consuming generated evidence may operate separately from the observer.
1569
+
1570
+ ## Browser and network safety
1571
+
1572
+ Running a browser introduces a security and privacy boundary that must be defined explicitly.
1573
+
1574
+ Before broad navigation support is implemented, the project must define behavior for matters such as:
1575
+
1576
+ - allowed URL schemes;
1577
+ - local versus remote targets;
1578
+ - redirects;
1579
+ - navigation timeouts;
1580
+ - certificate failures;
1581
+ - downloads;
1582
+ - popups;
1583
+ - browser permissions;
1584
+ - network requests;
1585
+ - unexpected navigation;
1586
+ - credential-bearing pages;
1587
+ - sensitive rendered data;
1588
+ - secret-bearing URLs or output;
1589
+ - cleanup of browser processes and temporary state.
1590
+
1591
+ The first version should remain intentionally conservative and local.
1592
+
1593
+ Safety behavior must be explicit rather than dependent on undocumented browser defaults.
1594
+
1595
+ External-reference support has a separate local-file privacy boundary. v0.7
1596
+ detects PNG/JPEG/WebP from header bytes, reads dimensions from bounded header
1597
+ bytes without decoding pixels, enforces file-size/dimension bounds, keeps
1598
+ operational paths out of semantic identity, and treats reference files as data,
1599
+ not executable instructions.
1600
+
1601
+ ## Target immutability
1602
+
1603
+ Observation is non-destructive by default.
1604
+
1605
+ The observer must not:
1606
+
1607
+ - edit the target application's files;
1608
+ - modify target source code;
1609
+ - commit target changes;
1610
+ - install dependencies into the target;
1611
+ - alter target configuration;
1612
+ - persist unintended application state;
1613
+ - perform destructive interactions merely to collect layout evidence.
1614
+
1615
+ The observed project is a target, not part of the observer repository.
1616
+
1617
+ Interactions such as:
1618
+
1619
+ - navigation;
1620
+ - viewport resize;
1621
+ - scrolling;
1622
+ - explicitly approved safe controls;
1623
+
1624
+ are acceptable when they are part of a defined observation scenario.
1625
+
1626
+ A coding agent or another external tool performs source changes.
1627
+
1628
+ Importing or evaluating an external reference does not authorize observer source edits or changes to the target application.
1629
+
1630
+ ## Preferred platform
1631
+
1632
+ Primary development platform:
1633
+
1634
+ - desktop developer workstation;
1635
+ - Windows first-class.
1636
+
1637
+ The implementation must avoid unnecessary Windows-specific assumptions.
1638
+
1639
+ Artifact paths, serialization, tests, and browser behavior should be designed so future ecosystem releases can satisfy the cross-platform validation expectations used by the broader `my-dev-kit` ecosystem.
1640
+
1641
+ Cross-platform screenshot byte identity should not be assumed unless explicitly established by testing.
1642
+
1643
+ Structured semantic evidence should remain the primary portable contract.
1644
+
1645
+ The same caution applies to reference/candidate image similarity: font rasterization, graphics environment, antialiasing, and browser/OS differences must not be treated as exact semantic equality unless explicitly proven.
1646
+
1647
+ ## Preferred implementation stack
1648
+
1649
+ Preferred language:
1650
+
1651
+ TypeScript.
1652
+
1653
+ Preferred runtime:
1654
+
1655
+ Node.js.
1656
+
1657
+ Preferred browser automation:
1658
+
1659
+ Playwright.
1660
+
1661
+ Initial public interface:
1662
+
1663
+ command-line interface (CLI).
1664
+
1665
+ The first scaffold should favor a TypeScript/Node.js command-line project rather than a web-application-first architecture.
1666
+
1667
+ The browser-observation engine must remain independent of command-line formatting so it can later support:
1668
+
1669
+ - command-line use;
1670
+ - programmatic use;
1671
+ - graphical local viewing;
1672
+ - automated regression workflows;
1673
+ - ecosystem adapters.
1674
+
1675
+ A later interactive viewer may use React or another suitable web user-interface stack.
1676
+
1677
+ Do not put Playwright/browser-control logic directly inside React presentation components.
1678
+
1679
+ Do not put canonical reference-evaluation logic directly inside viewer presentation components either.
1680
+
1681
+ Do not promise a stable public programmatic application programming interface merely because internal modules are reusable.
1682
+
1683
+ A public programmatic interface should become a compatibility commitment only when explicitly designed and tested.
1684
+
1685
+ ## Architectural direction
1686
+
1687
+ Use explicit ownership boundaries.
1688
+
1689
+ The smallest expected conceptual separation is:
1690
+
1691
+ ```text
1692
+ command-line interface
1693
+ ↓
1694
+ observation application/engine
1695
+ ↓
1696
+ browser adapter
1697
+ ↓
1698
+ runtime evidence
1699
+
1700
+ observation domain/schema
1701
+ ↓
1702
+ artifact writer
1703
+
1704
+ deterministic fixture infrastructure
1705
+ ↓
1706
+ browser-level validation
1707
+ ```
1708
+
1709
+ The architecture has added, and later capabilities may continue to add:
1710
+
1711
+ ```text
1712
+ relationship engine
1713
+ comparison engine
1714
+ contract engine
1715
+ bounded agent-context and correlation/export boundary
1716
+ coding-agent review workflow
1717
+ external visual-reference artifact/identity boundary
1718
+ reference-to-runtime binding
1719
+ reference/candidate structured evaluation
1720
+ viewer
1721
+ annotation system
1722
+ ```
1723
+
1724
+ These should extend the existing evidence model rather than creating parallel implementations.
1725
+
1726
+ Avoid speculative abstraction.
1727
+
1728
+ Do not create:
1729
+
1730
+ - a generic plugin framework without multiple real implementations;
1731
+ - a second observation engine for the viewer;
1732
+ - a second comparison implementation for the user interface;
1733
+ - a second contract engine for automated tests;
1734
+ - a viewer-only reference model or reference-evaluation engine;
1735
+ - annotation-only change semantics;
1736
+ - a generic ecosystem evidence framework before concrete integration requires one.
1737
+
1738
+ ## Initial product interface
1739
+
1740
+ The initial public interface is CLI-first.
1741
+
1742
+ The developer should be able to provide:
1743
+
1744
+ ```text
1745
+ target URL
1746
+ viewport
1747
+ observation targets
1748
+ output location
1749
+ ```
1750
+
1751
+ and receive:
1752
+
1753
+ ```text
1754
+ screenshot
1755
+ structured page observation
1756
+ structured target observations
1757
+ versioned observation artifact
1758
+ concise execution/result summary
1759
+ ```
1760
+
1761
+ The CLI should be suitable for both human and machine invocation.
1762
+
1763
+ Its architecture should leave room for:
1764
+
1765
+ - machine-readable output;
1766
+ - stable diagnostic codes;
1767
+ - explicit exit behavior;
1768
+ - separation between parseable output and human progress/diagnostics.
1769
+
1770
+ The CLI must not own browser logic directly.
1771
+
1772
+ The bounded agent context, static/runtime integration, text-driven coding-agent review, and non-graphical external-reference evidence foundation are on the core path after comparison/contracts. The graphical viewer and annotation system follow as human-interface enhancements.
1773
+
1774
+ ## Evidence boundedness
1775
+
1776
+ The observer must avoid collecting enormous amounts of runtime information merely because the browser exposes it.
1777
+
1778
+ Initial observation should be explicitly scoped.
1779
+
1780
+ Prefer:
1781
+
1782
+ ```text
1783
+ explicit observation targets
1784
+ + required page facts
1785
+ + required target facts
1786
+ ```
1787
+
1788
+ over:
1789
+
1790
+ ```text
1791
+ entire DOM
1792
+ + every style property
1793
+ + complete accessibility tree
1794
+ ```
1795
+
1796
+ Where evidence is bounded or truncated, the result should make the omission visible.
1797
+
1798
+ A bounded collection should expose enough information to distinguish:
1799
+
1800
+ ```text
1801
+ nothing existed
1802
+ ```
1803
+
1804
+ from:
1805
+
1806
+ ```text
1807
+ evidence existed but was omitted because of a limit
1808
+ ```
1809
+
1810
+ Required evidence adequacy must not mean merely that some evidence was captured.
1811
+
1812
+ If required configured evidence is missing, partial, or unavailable, the observer must say so.
1813
+
1814
+ The same rule applies to references. Do not send every region, pixel delta, style sample, or image byte to a coding agent when only a bounded subset is relevant to the requested correction. Reference artifacts and bounded fidelity context expose omission/truncation where limits matter.
1815
+
1816
+ ## Evidence provenance
1817
+
1818
+ Runtime evidence should remain traceable to its source.
1819
+
1820
+ Observation artifacts should record appropriate provenance such as:
1821
+
1822
+ - observer package version;
1823
+ - artifact schema version;
1824
+ - browser engine;
1825
+ - browser version;
1826
+ - target URL;
1827
+ - final URL;
1828
+ - viewport;
1829
+ - observation configuration;
1830
+ - target identity and locator;
1831
+ - observation method;
1832
+ - artifact references;
1833
+ - diagnostics;
1834
+ - limits and omissions;
1835
+ - capture identity;
1836
+ - derivation method for derived facts.
1837
+
1838
+ Reference evidence likewise preserves, as applicable:
1839
+
1840
+ - exact reference identity/version;
1841
+ - image reference and format/dimensions;
1842
+ - region identity and coordinate semantics;
1843
+ - authored requirements versus derived relationships;
1844
+ - applicability state;
1845
+ - approval/supersession state;
1846
+ - binding evidence;
1847
+ - tolerance/evaluation policy;
1848
+ - candidate observation identity;
1849
+ - diagnostics, limits, and omissions.
1850
+
1851
+ Naturally unstable metadata such as capture time should not become the only logical identity of an observation or reference.
1852
+
1853
+ ## Diagnostic behavior
1854
+
1855
+ Observation failure and partial evidence must be explainable.
1856
+
1857
+ The project should establish stable machine-readable diagnostics for cases such as:
1858
+
1859
+ - invalid request;
1860
+ - unsupported configuration;
1861
+ - navigation failure;
1862
+ - missing target;
1863
+ - ambiguous target;
1864
+ - hidden target;
1865
+ - unavailable browser evidence;
1866
+ - bounded/truncated evidence;
1867
+ - artifact write failure;
1868
+ - browser failure.
1869
+
1870
+ Current reference support likewise makes malformed/unsupported reference data,
1871
+ ambiguous/unavailable reference-to-runtime binding, incompatible
1872
+ reference/candidate state, unavailable candidate evidence, and bounded fidelity
1873
+ omissions explicit rather than fabricating normal values.
1874
+
1875
+ Do not silently select an arbitrary target when selection is ambiguous.
1876
+
1877
+ Do not represent unavailable evidence as a normal false or zero value.
1878
+
1879
+ Warnings, partial observations, invalid requests, and fatal failures must remain distinguishable.
1880
+
1881
+ ## Testing expectations
1882
+
1883
+ Testing is a core requirement.
1884
+
1885
+ The project should progressively include:
1886
+
1887
+ ```text
1888
+ unit tests
1889
+ → schema/serialization tests
1890
+ → browser adapter integration tests
1891
+ → deterministic browser fixture tests
1892
+ → comparison tests
1893
+ → contract tests
1894
+ → bounded agent-context and correlation tests
1895
+ → ecosystem compatibility fixtures
1896
+ → text-driven coding-agent workflow tests
1897
+ → external-reference artifact/identity/binding tests
1898
+ → reference applicability and structured fidelity tests
1899
+ → reference-driven coding-agent correction tests
1900
+ → viewer tests
1901
+ → annotation tests for runtime and reference contexts
1902
+ → full visual workflow tests
1903
+ ```
1904
+
1905
+ Important deterministic fixture scenarios should eventually include:
1906
+
1907
+ - normal desktop layout;
1908
+ - narrow navigation;
1909
+ - clipped navigation contents;
1910
+ - horizontal page overflow;
1911
+ - nested scrolling;
1912
+ - document scrolling;
1913
+ - footer after workspace;
1914
+ - overlapping regions;
1915
+ - mobile layout;
1916
+ - hidden elements;
1917
+ - expected dependent resizing;
1918
+ - protected-region regression;
1919
+ - external reference whose candidate has measurable geometry/spacing mismatch;
1920
+ - external reference with wrong theme/application-state candidate producing incompatibility;
1921
+ - asset-sensitive reference region;
1922
+ - reference-fidelity success coexisting with a protected-contract failure.
1923
+
1924
+ The first version should use controlled local fixture pages rather than depending on public internet pages for canonical test evidence.
1925
+
1926
+ Tests must distinguish:
1927
+
1928
+ ```text
1929
+ direct browser observation
1930
+ derived interpretation
1931
+ authored reference requirement
1932
+ ```
1933
+
1934
+ Screenshot evidence should not be treated as the only source of truth.
1935
+
1936
+ Cross-platform tests should distinguish semantic/layout evidence from rendering differences that may legitimately vary by operating system, browser build, fonts, or graphics environment.
1937
+
1938
+ ## Validation expectations
1939
+
1940
+ The project should maintain a trustworthy validation chain appropriate to its current capabilities.
1941
+
1942
+ At minimum, once established:
1943
+
1944
+ ```text
1945
+ typecheck
1946
+ lint
1947
+ unit/integration tests
1948
+ browser fixture tests
1949
+ production build when a graphical interface exists
1950
+ documentation checks when implemented
1951
+ ```
1952
+
1953
+ Browser-related functionality must always have browser-level evidence.
1954
+
1955
+ Passing static TypeScript validation alone is not sufficient for a browser-observation feature or a reference-driven workflow whose candidate side is browser-rendered.
1956
+
1957
+ Later ecosystem releases should also satisfy the coordinated compatibility and cross-platform validation expectations of the `my-dev-kit` ecosystem.
1958
+
1959
+ ## Performance expectations
1960
+
1961
+ The tool is a developer utility.
1962
+
1963
+ Correctness, boundedness, determinism, and inspectability are more important than extreme runtime optimization.
1964
+
1965
+ However:
1966
+
1967
+ - do not capture the entire Document Object Model when targeted evidence is sufficient;
1968
+ - do not emit enormous computed-style dumps;
1969
+ - do not take unnecessary screenshots;
1970
+ - do not repeatedly decode/copy the same reference image into every downstream artifact;
1971
+ - do not keep browser processes alive indefinitely;
1972
+ - make observation and reference scope explicit;
1973
+ - preserve evidence needed to explain conclusions;
1974
+ - avoid duplicating unchanged evidence unnecessarily.
1975
+
1976
+ ## Accessibility evidence
1977
+
1978
+ Where the browser exposes it reliably, capture useful semantic/accessibility information such as:
1979
+
1980
+ - role;
1981
+ - accessible name;
1982
+ - landmark identity;
1983
+ - relevant state.
1984
+
1985
+ This can help a human or LLM identify regions more reliably than position alone. Reference-to-runtime binding remains explicit and is never inferred merely because semantic evidence looks similar.
1986
+
1987
+ The project is not initially intended to replace a dedicated accessibility-audit product.
1988
+
1989
+ ## Post-v0.10 ecosystem evidence extensions
1990
+
1991
+ The initial project exclusions above describe what was not required to establish the original product. They do not prohibit later bounded capabilities when those capabilities preserve the same evidence ownership and local-first safety model.
1992
+
1993
+ After the full visual workflow, the product may add four explicitly bounded evidence domains:
1994
+
1995
+ ### Runtime diagnostics
1996
+
1997
+ Capture a bounded/redacted observation window for browser console messages, page errors, request transport failures, and HTTP response outcomes. Network transport failure must remain distinct from an HTTP response whose status represents an application/server error. Diagnostics remain evidence; they do not become an unrestricted network recorder, full HAR capture, or telemetry platform.
1998
+
1999
+ ### Controlled project state
2000
+
2001
+ Support project-supplied, explicitly governed local test-state/session setup with achieved-state evidence. State setup may include a project-provided authentication fixture or session artifact, but the observer does not become an identity provider, credential manager, general journey runner, or production authentication system. Secrets and credentials must not be persisted into ordinary evidence.
2002
+
2003
+ ### Local performance evidence
2004
+
2005
+ Capture bounded local performance measurements with explicit browser/environment provenance, observation windows, samples, and comparability rules. Missing or unsupported metrics remain unavailable. Performance evidence must not become an unexplained composite score or a claim that unlike environments are directly comparable.
2006
+
2007
+ ### Bounded browser and viewport matrices
2008
+
2009
+ Allow a deliberately bounded set of browser-engine/viewport combinations when the adapter architecture and fixtures support them. Every engine retains its identity and capability differences. Adding selected Firefox or WebKit evidence later does not turn the product into a complete cross-browser testing service, and evidence from incompatible engines must not be merged as though it were one baseline.
2010
+
2011
+ These domains remain subordinate to the core principles: explicit scope, bounded evidence, deterministic artifacts where practical, inspectability, separation of observed and derived claims, and no source editing.
2012
+
2013
+ ## Inspectability
2014
+
2015
+ Observation and regression results must be explainable.
2016
+
2017
+ A useful result should identify:
2018
+
2019
+ ```text
2020
+ what was observed or explicitly referenced
2021
+ where it was observed/referenced
2022
+ what changed or still differs
2023
+ before/reference value
2024
+ after/candidate value
2025
+ difference
2026
+ expected condition
2027
+ actual condition
2028
+ contract, reference requirement, or relationship involved
2029
+ supporting artifact
2030
+ supporting screenshot/reference
2031
+ ```
2032
+
2033
+ Avoid unexplained scores.
2034
+
2035
+ Avoid opaque artificial-intelligence classification in the core validation path.
2036
+
2037
+ An LLM may reason over the evidence, but the evidence producer itself should remain inspectable.
2038
+
2039
+ ## Determinism
2040
+
2041
+ Given:
2042
+
2043
+ - the same target build;
2044
+ - the same browser version;
2045
+ - the same viewport;
2046
+ - the same observation configuration;
2047
+ - the same deterministic fixture state;
2048
+
2049
+ the structured observation should be stable enough for meaningful comparison.
2050
+
2051
+ Given the same approved reference content, reference configuration, region definitions, applicability identity, and tolerance policy, the reference's logical identity and fidelity evaluation are deterministic according to the v0.7 contract.
2052
+
2053
+ Fields that are naturally unstable must either:
2054
+
2055
+ - be normalized;
2056
+ - be excluded from logical comparison;
2057
+ - or be explicitly identified as unstable metadata.
2058
+
2059
+ Deterministic target ordering, reference-region ordering, diagnostic ordering, serialization, and artifact references should be preferred where practical.
2060
+
2061
+ ## Non-goals for the initial project
2062
+
2063
+ The initial project is not:
2064
+
2065
+ - a replacement for browser developer tools;
2066
+ - a replacement for Playwright;
2067
+ - a replacement for `my-dev-kit`;
2068
+ - a replacement for `my-dev-kit-orchestrator`;
2069
+ - a replacement for `my-dev-kit-lab`;
2070
+ - an autonomous frontend designer;
2071
+ - an autonomous coding agent;
2072
+ - a visual website builder;
2073
+ - a hosted screenshot service;
2074
+ - a cloud browser farm;
2075
+ - a full accessibility scanner;
2076
+ - a complete cross-browser testing service;
2077
+ - a pixel-perfect visual-diff-only system;
2078
+ - a Figma or Canva replacement;
2079
+ - a screenshot-cloning SaaS;
2080
+ - an autonomous raster-to-HTML/CSS generator;
2081
+ - an automatic logo/vector reconstruction system;
2082
+ - a general-purpose computer-vision framework;
2083
+ - a source-code editor;
2084
+ - a deployment system.
2085
+
2086
+ The initial project does not need:
2087
+
2088
+ - authentication;
2089
+ - payments;
2090
+ - advertising;
2091
+ - multi-user collaboration;
2092
+ - cloud persistence;
2093
+ - remote browser infrastructure;
2094
+ - external LLM APIs;
2095
+ - production hosting;
2096
+ - Firefox or WebKit support;
2097
+ - source ownership;
2098
+ - static repository indexing;
2099
+ - orchestrator integration;
2100
+ - lab integration;
2101
+ - external visual-reference evaluation;
2102
+ - visual annotation;
2103
+ - comparison;
2104
+ - regression contracts.
2105
+
2106
+ Those capabilities may appear later according to Project Milestones and `ROADMAP.md`.
2107
+
2108
+ ## Explicit product principles
2109
+
2110
+ 1. Observe before inferring.
2111
+ 2. Browser runtime is authoritative for rendered geometry.
2112
+ 3. Source code and rendered output are different evidence domains.
2113
+ 4. External desired-design references are a third evidence domain, distinct from runtime observations and source code.
2114
+ 5. `my-dev-kit` owns static repository/source evidence; `my-frontend-observer` owns runtime browser evidence and its structured reference-evidence boundary.
2115
+ 6. Stable runtime-region identity does not automatically imply known source ownership.
2116
+ 7. Reference-region identity does not automatically imply runtime-target identity or source ownership.
2117
+ 8. Observed dimensions are measurements, not automatically fixed design constants.
2118
+ 9. A visible reference pixel is not automatically a hard requirement.
2119
+ 10. Prefer relationship-based layout requirements when they better represent user intent.
2120
+ 11. Distinguish direct browser facts, direct image measurements, authored requirements, and derived interpretations.
2121
+ 12. Preserve raw evidence behind normalized and summarized evidence.
2122
+ 13. Keep evidence bounded and make omissions explicit.
2123
+ 14. Never represent unavailable evidence as if it were an observed false or zero.
2124
+ 15. Never claim a visual requirement passed solely because a styling declaration looks correct.
2125
+ 16. Never claim reference fidelity from pixel similarity alone when structured evidence is available or required.
2126
+ 17. A requested change may legitimately cause dependent changes.
2127
+ 18. Distinguish requested changes, expected dependent changes, protected properties, preserved invariants, and unexpected changes.
2128
+ 19. A local requested change or reference match does not authorize unrelated rendered changes.
2129
+ 20. Previously approved frontend invariants remain active unless the user explicitly supersedes them.
2130
+ 21. Imported reference, approved reference, baseline approval, reference supersession, and baseline supersession are separate states/actions.
2131
+ 22. Make regressions and fidelity failures explainable.
2132
+ 23. Keep observation and reference evaluation non-destructive.
2133
+ 24. Keep artifacts local-first, versioned, portable, and inspectable.
2134
+ 25. Separate browser observation from static source analysis.
2135
+ 26. Separate evidence production from workflow orchestration and downstream evaluation.
2136
+ 27. Human visual intent must eventually be representable alongside machine measurements and external desired-design references.
2137
+ 28. Deep ecosystem integration should use explicit contracts and adapters rather than duplicated responsibilities.
2138
+ 29. Do not introduce speculative cross-project coupling before a real consumer requires it.
2139
+ 30. The viewer and annotation layers consume canonical reference/comparison/contract evidence; they do not redefine it.
2140
+
2141
+ ## Documentation and planning principles
2142
+
2143
+ Documentation must distinguish current implemented behavior from future intended behavior.
2144
+
2145
+ Current-state documentation should accurately record what exists.
2146
+
2147
+ Forward-looking planning documents should preserve enough local design context for future LLM planning without requiring critical intent to be reconstructed from many unrelated bookkeeping documents.
2148
+
2149
+ In particular:
2150
+
2151
+ ```text
2152
+ Project Description
2153
+ → durable product intent
2154
+ → responsibility boundaries
2155
+ → long-term capability model
2156
+
2157
+ Project Milestones
2158
+ → ordered capability development
2159
+ → major requirements
2160
+ → acceptance expectations
2161
+ → cross-milestone invariants
2162
+
2163
+ ROADMAP.md
2164
+ → high-level version specifications
2165
+ → version goals
2166
+ → required capabilities
2167
+ → architectural constraints
2168
+ → dependencies
2169
+ → exclusions
2170
+ → acceptance expectations
2171
+ ```
2172
+
2173
+ `ROADMAP.md` must not predefine implementation batches.
2174
+
2175
+ When implementation of a roadmap version begins, the planner should:
2176
+
2177
+ ```text
2178
+ read the roadmap version
2179
+ → inspect current repository state
2180
+ → obtain required architecture/retrieval evidence
2181
+ → design the implementation steps
2182
+ → divide those steps into appropriate implementation batches
2183
+ → execute and validate those batches
2184
+ ```
2185
+
2186
+ Forward-looking requirements may intentionally appear in more than one planning document when doing so prevents future planning context from becoming fragmented.
2187
+
2188
+ ## Long-term product direction
2189
+
2190
+ The long-term goal is to create a reliable communication and validation bridge between:
2191
+
2192
+ ```text
2193
+ human visual intent
2194
+ approved external desired-design references where applicable
2195
+ rendered frontend reality
2196
+ static repository evidence
2197
+ LLM reasoning
2198
+ coding-agent implementation
2199
+ ```
2200
+
2201
+ The critical path has established:
2202
+
2203
+ ```text
2204
+ render and observe
2205
+ → identify stable regions and runtime behavior
2206
+ → compare
2207
+ → enforce requested/dependent/protected/preserved scope
2208
+ → combine bounded runtime and static evidence
2209
+ → establish external-reference identity/regions/requirements/applicability/binding
2210
+ → evaluate reference fidelity
2211
+ → provide bounded context to an external coding agent
2212
+ → rerender and reject fidelity or protected-contract regressions
2213
+ ```
2214
+
2215
+ Only after those core workflows work should the human visual branch add:
2216
+
2217
+ ```text
2218
+ viewer with reference/candidate inspection
2219
+ → structured annotation on runtime or reference images
2220
+ → full visual human–LLM workflow with both entry modes
2221
+ ```
2222
+
2223
+ The desired eventual visual cycle is:
2224
+
2225
+ ```text
2226
+ render/current candidate
2227
+ + optional approved external reference
2228
+ → observe
2229
+ → identify stable runtime regions and reference regions
2230
+ → measure geometry and behavior
2231
+ → evaluate reference applicability/fidelity where applicable
2232
+ → show human
2233
+ → annotate/request change on runtime or reference
2234
+ → define requested/dependent/protected/preserved scope
2235
+ → combine bounded runtime/reference and static evidence
2236
+ → provide context to LLM
2237
+ → coding agent implements
2238
+ → rerender
2239
+ → compare before/after
2240
+ → reevaluate reference/candidate fidelity
2241
+ → rerun preserved contracts
2242
+ → identify unexpected changes
2243
+ → approve or correct
2244
+ → establish new baseline and/or explicitly supersede reference according to policy
2245
+ → repeat
2246
+ ```
2247
+
2248
+ The project succeeds when an LLM no longer needs to guess what a frontend looks like from source code alone, when a human can communicate visual intent or an approved desired design without translating every design idea into implementation terminology, and when a frontend change cannot be considered successful while silently breaking previously approved rendered behavior.