@dailephd/my-frontend-observer 0.9.1 → 0.10.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +490 -471
- package/LICENSE +21 -21
- package/README.md +375 -357
- package/dist/application/projectCheckService.d.ts +6 -0
- package/dist/application/projectCheckService.js +8 -1
- package/dist/application/projectCheckService.js.map +1 -1
- package/dist/application/projectWorkflowService.d.ts +7 -2
- package/dist/application/projectWorkflowService.js +10 -3
- package/dist/application/projectWorkflowService.js.map +1 -1
- package/dist/application/visualChangeAgentHandoffService.d.ts +28 -0
- package/dist/application/visualChangeAgentHandoffService.js +111 -0
- package/dist/application/visualChangeAgentHandoffService.js.map +1 -0
- package/dist/application/visualChangeProjectWorkflowService.d.ts +95 -0
- package/dist/application/visualChangeProjectWorkflowService.js +376 -0
- package/dist/application/visualChangeProjectWorkflowService.js.map +1 -0
- package/dist/application/visualChangeReviewService.d.ts +50 -0
- package/dist/application/visualChangeReviewService.js +69 -0
- package/dist/application/visualChangeReviewService.js.map +1 -0
- package/dist/application/visualChangeWorkflowPersistenceService.d.ts +26 -0
- package/dist/application/visualChangeWorkflowPersistenceService.js +15 -0
- package/dist/application/visualChangeWorkflowPersistenceService.js.map +1 -0
- package/dist/artifacts/visualChangeWorkflowArtifactReader.d.ts +9 -0
- package/dist/artifacts/visualChangeWorkflowArtifactReader.js +47 -0
- package/dist/artifacts/visualChangeWorkflowArtifactReader.js.map +1 -0
- package/dist/artifacts/visualChangeWorkflowArtifactWriter.d.ts +20 -0
- package/dist/artifacts/visualChangeWorkflowArtifactWriter.js +41 -0
- package/dist/artifacts/visualChangeWorkflowArtifactWriter.js.map +1 -0
- package/dist/cli.js +9 -7
- package/dist/cli.js.map +1 -1
- package/dist/domain/visualChangeAgentHandoff.d.ts +82 -0
- package/dist/domain/visualChangeAgentHandoff.js +80 -0
- package/dist/domain/visualChangeAgentHandoff.js.map +1 -0
- package/dist/domain/visualChangeAgentHandoffSerialization.d.ts +2 -0
- package/dist/domain/visualChangeAgentHandoffSerialization.js +11 -0
- package/dist/domain/visualChangeAgentHandoffSerialization.js.map +1 -0
- package/dist/domain/visualChangeCycle.d.ts +8 -0
- package/dist/domain/visualChangeCycle.js +7 -0
- package/dist/domain/visualChangeCycle.js.map +1 -0
- package/dist/domain/visualChangeWorkflow.d.ts +125 -0
- package/dist/domain/visualChangeWorkflow.js +109 -0
- package/dist/domain/visualChangeWorkflow.js.map +1 -0
- package/dist/domain/visualChangeWorkflowIdentity.d.ts +5 -0
- package/dist/domain/visualChangeWorkflowIdentity.js +24 -0
- package/dist/domain/visualChangeWorkflowIdentity.js.map +1 -0
- package/dist/index.d.ts +21 -1
- package/dist/index.js +12 -1
- package/dist/index.js.map +1 -1
- package/dist/projectWorkflow/projectPaths.d.ts +3 -0
- package/dist/projectWorkflow/projectPaths.js +7 -0
- package/dist/projectWorkflow/projectPaths.js.map +1 -1
- package/dist/viewer/assets/index-DglJ6f28.css +1 -0
- package/dist/viewer/assets/index-DsODREY5.js +9 -0
- package/dist/viewer/index.html +15 -15
- package/dist/viewer/sw.js +1 -1
- package/dist/viewerServer/evidence/classify.d.ts +3 -1
- package/dist/viewerServer/evidence/classify.js +10 -0
- package/dist/viewerServer/evidence/classify.js.map +1 -1
- package/dist/viewerServer/evidence/handles.js +1 -0
- package/dist/viewerServer/evidence/handles.js.map +1 -1
- package/dist/viewerServer/evidence/projection.d.ts +5 -0
- package/dist/viewerServer/evidence/projection.js +19 -0
- package/dist/viewerServer/evidence/projection.js.map +1 -1
- package/dist/viewerServer/evidence/visualChangeWorkflowView.d.ts +31 -0
- package/dist/viewerServer/evidence/visualChangeWorkflowView.js +36 -0
- package/dist/viewerServer/evidence/visualChangeWorkflowView.js.map +1 -0
- package/dist/viewerServer/httpServer.js +323 -1
- package/dist/viewerServer/httpServer.js.map +1 -1
- package/dist/viewerServer/referenceApproval.d.ts +22 -0
- package/dist/viewerServer/referenceApproval.js +42 -0
- package/dist/viewerServer/referenceApproval.js.map +1 -0
- package/dist/viewerServer/referenceVisualChangeAuthoring.d.ts +28 -0
- package/dist/viewerServer/referenceVisualChangeAuthoring.js +134 -0
- package/dist/viewerServer/referenceVisualChangeAuthoring.js.map +1 -0
- package/dist/viewerServer/runtimeVisualChangeAuthoring.d.ts +33 -0
- package/dist/viewerServer/runtimeVisualChangeAuthoring.js +81 -0
- package/dist/viewerServer/runtimeVisualChangeAuthoring.js.map +1 -0
- package/dist/viewerServer/visualChangeAuthoring.d.ts +46 -0
- package/dist/viewerServer/visualChangeAuthoring.js +63 -0
- package/dist/viewerServer/visualChangeAuthoring.js.map +1 -0
- package/dist/viewerServer/visualChangeHandoff.d.ts +23 -0
- package/dist/viewerServer/visualChangeHandoff.js +31 -0
- package/dist/viewerServer/visualChangeHandoff.js.map +1 -0
- package/dist/viewerServer/visualChangeReview.d.ts +30 -0
- package/dist/viewerServer/visualChangeReview.js +46 -0
- package/dist/viewerServer/visualChangeReview.js.map +1 -0
- package/docs/ARCHITECTURE.md +1394 -1373
- package/docs/CI_CD.md +349 -327
- package/docs/COMMANDS.md +1035 -1012
- package/docs/CONTRACTS.md +1971 -1926
- package/docs/CURRENT_STATE.md +1277 -1238
- package/docs/DEVELOPMENT.md +240 -237
- package/docs/DOCUMENTATION_PRESERVATION_POLICY.md +50 -50
- package/docs/PROJECT_DESCRIPTION.md +2248 -2224
- package/docs/PROJECT_MILESTONES.md +2681 -2558
- package/docs/PROJECT_OVERVIEW.md +200 -191
- package/docs/QUICKSTART.md +100 -96
- package/docs/RELEASE.md +37 -33
- package/docs/ROADMAP.md +1105 -1033
- package/docs/SECURITY.md +297 -275
- package/docs/WORKFLOWS.md +806 -770
- package/docs/plans/v0.10-implementation-plan.md +1509 -0
- package/docs/plans/v0.8-implementation-plan.md +655 -655
- package/docs/plans/v0.8.1-cli-usability-patch-plan.md +505 -505
- package/docs/plans/v0.9-implementation-plan.md +1529 -1529
- package/docs/plans/v0.9.1-implementation-plan.md +468 -468
- package/docs/reports/v0.10-batch1-visual-change-workflow-foundation.md +102 -0
- package/docs/reports/v0.10-batch2-project-composition-check-recording.md +103 -0
- package/docs/reports/v0.10-batch3-viewer-visual-change-workspace.md +93 -0
- package/docs/reports/v0.10-batch4-actual-frontend-entry.md +59 -0
- package/docs/reports/v0.10-batch5-reference-driven-entry.md +238 -0
- package/docs/reports/v0.10-batch6-coding-agent-handoff.md +85 -0
- package/docs/reports/v0.10-batch7-correction-review-acceptance.md +145 -0
- package/docs/reports/v0.10-batch8-integrated-acceptance.md +109 -0
- package/docs/reports/v0.10-implementation-completeness-documentation-reconciliation.md +344 -0
- package/docs/reports/v0.10-pre-release-readiness.md +120 -0
- package/docs/reports/v0.10-release-preparation.md +70 -0
- package/docs/reports/v0.10.1-project-check-baseline-context-implementation.md +86 -0
- package/docs/reports/v0.7-bounded-fidelity-context-prompt7.md +243 -243
- package/docs/reports/v0.7-implementation-completeness-documentation-reconciliation.md +497 -497
- package/docs/reports/v0.7-pre-release-readiness.md +337 -337
- package/docs/reports/v0.7-reference-binding-prompt5.md +223 -223
- package/docs/reports/v0.7-reference-compatibility-prompt4.md +234 -234
- package/docs/reports/v0.7-reference-correction-workflow-prompt8.md +222 -222
- package/docs/reports/v0.7-reference-fidelity-prompt6.md +216 -216
- package/docs/reports/v0.7-reference-foundation-prompt1.md +151 -151
- package/docs/reports/v0.7-reference-regions-prompt2.md +195 -195
- package/docs/reports/v0.7-reference-requirements-prompt3.md +217 -217
- package/docs/reports/v0.7-release-prep.md +423 -423
- package/docs/reports/v0.8-binding-fidelity-interaction-batch6.md +279 -279
- package/docs/reports/v0.8-bounded-context-correlation-batch7.md +233 -233
- package/docs/reports/v0.8-comparison-contract-inspection-batch4.md +279 -279
- package/docs/reports/v0.8-evidence-index-readers-batch2.md +247 -247
- package/docs/reports/v0.8-implementation-completeness-documentation-reconciliation.md +741 -741
- package/docs/reports/v0.8-integrated-viewer-acceptance-batch8.md +128 -128
- package/docs/reports/v0.8-observation-svg-inspection-batch3.md +223 -223
- package/docs/reports/v0.8-prerelease-readiness-cross-platform-security-code-rot.md +687 -687
- package/docs/reports/v0.8-reference-candidate-inspection-batch5.md +232 -232
- package/docs/reports/v0.8-viewer-runtime-pwa-batch1.md +278 -278
- package/docs/reports/v0.8.1-implementation-completeness-documentation-reconciliation.md +114 -114
- package/docs/reports/v0.8.1-prerelease-readiness-cross-platform-security-code-rot.md +170 -170
- package/docs/reports/v0.9-architecture-retrieval.md +14 -37
- package/docs/reports/v0.9-final-pre-release-readiness.md +209 -209
- package/docs/reports/v0.9-final-readiness-corrections.md +530 -530
- package/docs/reports/v0.9-pre-release-readiness.md +169 -169
- package/docs/reports/v0.9.1-batch1-pwa-hard-gate-isolation.md +359 -359
- package/docs/reports/v0.9.1-batch2-hard-gate-validation-integration.md +262 -262
- package/docs/reports/v0.9.1-pre-release-readiness.md +206 -206
- package/package.json +59 -59
- package/dist/viewer/assets/index-BN41MI7m.css +0 -1
- package/dist/viewer/assets/index-CkKXnlrI.js +0 -9
|
@@ -1,216 +1,216 @@
|
|
|
1
|
-
# v0.7 Prompt 6 — Structured Reference-vs-Candidate Fidelity Evaluation
|
|
2
|
-
|
|
3
|
-
**VERDICT: PASS_V0_7_REFERENCE_FIDELITY_PROMPT6**
|
|
4
|
-
|
|
5
|
-
## Repository / branch / heads
|
|
6
|
-
|
|
7
|
-
- Repository: `my-frontend-observer` (path: `Z:\Users\newuser\Projects\my-frontend-observer`)
|
|
8
|
-
- Branch: `implementation/v0.7-reference-fidelity`, branched from the exact completed Prompt 5 HEAD
|
|
9
|
-
- Starting HEAD (branch point, Prompt 5 report commit): `74c710daa62cd5a56f1b1b2eac3a43d24dfe69fb`
|
|
10
|
-
- Prompt 5 base HEAD confirmed to contain both required commits: `74cc148` (implementation) and `74c710d` (report), and the full completed Prompt 1–4 lineage — verified via `git log --oneline` before branching.
|
|
11
|
-
- Implementation commit (this prompt): `91f5e736a1a8472984b0c8c7a7e179e307f1a5ce` — "Add v0.7 Prompt 6 structured reference-vs-candidate fidelity evaluation"
|
|
12
|
-
- Ending HEAD: the report commit that follows this file's commit.
|
|
13
|
-
|
|
14
|
-
## Git status
|
|
15
|
-
|
|
16
|
-
Preflight (`git status --short`) showed a clean working tree at the exact Prompt 5 report HEAD; `git stash list` showed exactly the one preserved Prompt 1 stray-fork-writes entry. The branch was created with `git checkout -b implementation/v0.7-reference-fidelity` and ancestry verified with `git merge-base --is-ancestor 74c710d HEAD`. No reset/clean/discard operation was used; the Prompt 1 stash was neither applied nor dropped. Post-implementation `git status --short` is clean except for this report file (staged and committed separately, per convention).
|
|
17
|
-
|
|
18
|
-
## Tooling
|
|
19
|
-
|
|
20
|
-
- Resolved latest published `@dailephd/my-dev-kit`: `1.12.3`, pinned exactly via `npx -y @dailephd/my-dev-kit@1.12.3`.
|
|
21
|
-
- Fresh Prompt 6 repository index built under a repository-local, gitignored root: `.my-dev-kit-context/index-prompt6/` (57 files, 868 symbols indexed) — not reused from Prompt 1–5's indexes.
|
|
22
|
-
|
|
23
|
-
## Prompt 1–5 contracts verified
|
|
24
|
-
|
|
25
|
-
Read directly (source, not memory) before writing any code:
|
|
26
|
-
|
|
27
|
-
- `src/domain/externalReferenceRequirements.ts` — the full Prompt 3 requirement/tolerance/expectation/adequacy model, including the exact `RELATIONSHIP_FAMILY_GROUPS`-scoped lookup fix and `deriveReferenceRequirementAdequacy`'s "100% unavailable requirements ⇒ inadequate" rule.
|
|
28
|
-
- `src/domain/externalReferenceRegions.ts` — `ReferenceRegionGeometry`'s exact field set (`x`/`y`/`width`/`height`/`right`/`bottom`/`centerX`/`centerY`), reused as the shared shape for both reference and (converted) candidate geometry.
|
|
29
|
-
- `src/domain/externalReference.ts` — `isImportedExternalReferenceArtifact`/`isApprovedExternalReferenceArtifact` type guards, reused to read image dimensions uniformly across both lifecycle variants without duplicating validation logic.
|
|
30
|
-
- `src/domain/externalReferenceCompatibility.ts` (Prompt 4) — `evaluateReferenceCandidateCompatibility`'s exact signature and result shape, reused verbatim as the compatibility gate.
|
|
31
|
-
- `src/domain/externalReferenceRuntimeBinding.ts` (Prompt 5) — `evaluateReferenceRuntimeBindings`'s exact signature/result shape and its own case-insensitive target-name-resolution convention (mirrored, not imported, per the established per-module small-helper convention).
|
|
32
|
-
- `src/domain/schema.ts` (v0.2) — `TargetGeometry`, `TargetVisibility`, `TargetResolution`/`TargetSelectionStatus`, and `ObservationArtifact.targetEvidence`'s exact-configured-name keying.
|
|
33
|
-
- `src/domain/comparisonEngine.ts` (v0.4) — `targetPresence` (already additively exported by Prompt 5), reused rather than re-derived.
|
|
34
|
-
- `src/domain/relationships.ts` (v0.4) — `deriveLayoutRelationships`, `PairwiseLayoutRelationship`'s `subjectTarget`/`relatedTarget` fields, `UnresolvedRelationshipTarget`/`UNRESOLVED_TARGET_REASONS` (confirming hidden targets are already excluded from relationship derivation by v0.4 itself), and every pairwise relationship-family constant.
|
|
35
|
-
- `src/domain/frontendContracts.ts`/`frontendContractEvaluation.ts` (v0.5) — `CLAUSE_RESULT_STATUSES`/`ClauseEvaluationResult` (the `pass`/`fail`/`unavailable`/`conflict` precedent), `evaluateFrontendContract`'s overall-verdict aggregation rule (`every clause pass && no unexpected changes ⇒ PASS`), `clauseMode`'s `required`/`permitted` handling (confirming it is a v0.5-only, directional-primitive-specific concept never mirrored in Prompt 3's own expectation/adequacy derivation), and `toleranceToPx`'s percent-denominator convention (`(amount/100) * abs(beforeValue)`), reused independently for reference-image-pixel units.
|
|
36
|
-
- `src/application/comparisonService.ts` — the exact `compareAndPersistFromArtifactRoots` application-layer shape (read-both-artifacts-then-delegate-to-the-pure-function), mirrored for `evaluateReferenceCandidateFidelityFromArtifactRoots`.
|
|
37
|
-
- `src/cli.ts` — the full `evaluate-contract` command (help text, argument parser, `--enforce` exit-code handling) as the direct precedent for `evaluate-reference-fidelity`, and every existing `--*-file` loader (`loadRegionsFile`/`loadRequirementsFile`/`loadApplicabilityFile`) as the precedent for `loadBindingsFile`.
|
|
38
|
-
|
|
39
|
-
No architecture blocker was hit; no `BLOCKED_ESCALATE_TO_FULL_STAGE_CONTEXT` condition applied at any point.
|
|
40
|
-
|
|
41
|
-
## Precedent review outcomes
|
|
42
|
-
|
|
43
|
-
- **A. Result vocabulary** — `pass`/`fail`/`unavailable` reused as the same honest three-state shape v0.5's `CLAUSE_RESULT_STATUSES` already established, but as an independently-owned constant (`REFERENCE_REQUIREMENT_FIDELITY_STATUSES`) deliberately excluding `'conflict'` — Prompt 6 has no cross-requirement authoring-conflict concept (each requirement is evaluated independently against its own subject).
|
|
44
|
-
- **B. Overall fidelity** — a compatible/evaluable pair reaches `state: 'pass'` only when every requirement result is `'pass'`; any `'fail'`/`'unavailable'` forces `'fail'`. An incompatible or reference-inadequate pair never reaches that computation at all — a distinct `not-evaluated` state, with `blockedBy: 'reference-inadequate' | 'incompatible'`, is returned instead.
|
|
45
|
-
- **C. Relationship evaluation** — reuses Prompt 2's `deriveReferenceRequirementExpectation` (reference side) and v0.4's `deriveLayoutRelationships` (candidate side) exactly; no duplicated geometry predicate exists anywhere in the new module.
|
|
46
|
-
- **D. Value derivation** — reference expected values always come from `deriveReferenceRequirementExpectation`/`deriveReferenceRequirementMeasurement` (Prompt 3), never recomputed by hand; candidate values always come from `ObservationArtifact.targetEvidence`'s existing `TargetGeometry`/`TargetVisibility` evidence, never a second derivation.
|
|
47
|
-
- **E. Persistence** — no new persisted artifact family; documented decision below.
|
|
48
|
-
|
|
49
|
-
## Fidelity evaluator type/function names
|
|
50
|
-
|
|
51
|
-
- `domain/externalReferenceFidelity.ts`: `evaluateReferenceCandidateFidelity(reference, candidate, bindingDeclarations, options?)` → `{ ok: true; evaluation: ReferenceCandidateFidelityEvaluation } | { ok: false; reason: string }`.
|
|
52
|
-
- `application/referenceFidelityEvaluationService.ts`: `evaluateReferenceCandidateFidelityFromArtifactRoots(referenceRoot, candidateRoot, bindingDeclarations, options?)` — the CLI-facing, artifact-root-reading counterpart.
|
|
53
|
-
- CLI: `evaluate-reference-fidelity`.
|
|
54
|
-
|
|
55
|
-
## Evaluation ordering / gates
|
|
56
|
-
|
|
57
|
-
Frozen order, exactly as specified and never reordered: reference structural validation (`isValidExternalReferenceArtifact`) → candidate structural validation (`isValidObservationArtifact`) → binding-declaration structural validation (`isValidReferenceRuntimeBindingDeclarations`, invoked inside the Prompt 5 evaluator) → Prompt 3 reference adequacy (`deriveReferenceRequirementAdequacy`) → Prompt 4 compatibility (`evaluateReferenceCandidateCompatibility`) → Prompt 5 binding evaluation (`evaluateReferenceRuntimeBindings`) → per-requirement candidate-evidence/coordinate-mapping checks → per-requirement tolerance/relationship comparison → overall result.
|
|
58
|
-
|
|
59
|
-
The first three (structural) failures return `{ ok: false, reason }` — a caller/config error, never a fidelity outcome. `adequacy.status === 'inadequate'` and `compatibility.state === 'incomparable'` each short-circuit to `state: 'not-evaluated'` with `requirementResults: []` — verified by dedicated tests (behaviors G/H). `adequacy.status === 'partial'` does **not** block evaluation; it proceeds normally, and the specific reference-side-unavailable requirements independently re-derive their own `unavailable` result at the per-requirement stage (not looked up from the adequacy result — verified by the "reference itself does not exhibit" test, which requires a second, evaluable requirement to keep adequacy at `partial` rather than `inadequate`, confirming the two computations are genuinely independent).
|
|
60
|
-
|
|
61
|
-
## Coordinate-mapping model
|
|
62
|
-
|
|
63
|
-
One explicit, deterministic full-frame scale: `scaleX = referenceImageWidth / applicableViewportWidth`, `scaleY = referenceImageHeight / applicableViewportHeight`, derived once per evaluation call from `reference.applicability.viewport` (Prompt 4) and the reference image's own pixel dimensions (Prompt 1, read uniformly across both lifecycle variants via the existing `isImportedExternalReferenceArtifact`/`isApprovedExternalReferenceArtifact` type guards). Candidate `TargetGeometry` (CSS pixels) is reshaped into a `ReferenceRegionGeometry`-shaped value (adding `centerX`/`centerY`, computed identically to `deriveReferenceRegionGeometry`) both raw and scaled; horizontal fields (`x`/`width`/`right`/`centerX`) always scale by `scaleX`, vertical fields (`y`/`height`/`bottom`/`centerY`) always scale by `scaleY` — this falls out automatically from the per-field conversion, never a hand-picked per-property axis table. `region-measurement` subjects reuse Prompt 3's exact `deriveReferenceRequirementMeasurement` on the converted geometries directly, rather than a parallel "runtime version" of the same gap/delta formulas.
|
|
64
|
-
|
|
65
|
-
## Image/viewport scale rule and aspect-ratio rule
|
|
66
|
-
|
|
67
|
-
`ASPECT_RATIO_MAPPING_TOLERANCE = 0.01` (1% relative difference between `scaleX` and `scaleY`) is an independently-owned, tiny coordinate-mapping-validity constant — never a user-authored Prompt 3 design tolerance. If `scaleX`/`scaleY` disagree beyond this bound, or `reference.applicability.viewport` is absent entirely, `deriveCoordinateScale` returns `{ ok: false, reason }` and every numeric (`region-property`/`region-measurement`) requirement in that evaluation becomes `unavailable`/`coordinate-mapping-unavailable` — verified by dedicated tests (behaviors E and F). No cropping, offset, rotation, or perspective registration is implemented anywhere; only this single bounded full-frame check. Categorical `region-relationship` requirements are unaffected by an unavailable scale (they never call `deriveCoordinateScale` at all).
|
|
68
|
-
|
|
69
|
-
## Numeric precision rule
|
|
70
|
-
|
|
71
|
-
No intermediate rounding anywhere in the arithmetic path (`referenceValue`, `candidateRawValue`, `candidateValue`, `delta` are all plain JS `number`s carried through at full floating-point precision). Tolerance comparison is a single `Math.abs(delta) <= allowed` check with no added epsilon — verified by the exact-boundary test (a delta equal to the tolerance passes; a delta a tiny fraction over it fails), matching the task's own worked boundary example (428 passes, 428.0001 fails against a reference of 424 ± 4).
|
|
72
|
-
|
|
73
|
-
## Requirement result vocabulary
|
|
74
|
-
|
|
75
|
-
`REFERENCE_REQUIREMENT_FIDELITY_STATUSES = ['pass', 'fail', 'unavailable']`. `reasonCode` (one of `reference-evidence-unavailable`, `reference-relationship-not-exhibited`, `binding-unavailable`, `candidate-evidence-unavailable`, `coordinate-mapping-unavailable`) and `detail` are present only when `status === 'unavailable'` — a `fail` is already fully explained by its numeric (`referenceValue`/`candidateRawValue`/`candidateValue`/`delta`/`tolerance`) or relationship (`expectedRelationship`/`actualRelationship`) fields, so no reason code is needed there.
|
|
76
|
-
|
|
77
|
-
## Overall fidelity vocabulary
|
|
78
|
-
|
|
79
|
-
`REFERENCE_FIDELITY_STATES = ['not-evaluated', 'pass', 'fail']`, with `REFERENCE_FIDELITY_BLOCK_REASONS = ['reference-inadequate', 'incompatible']` populated only when `state === 'not-evaluated'`. There is no separate "partial pass" bucket once evaluation has actually run: any `unavailable` requirement result forces `state: 'fail'`, satisfying "any required fail or unavailable must prevent overall PASS" (section 34) without inventing a fourth top-level state — an `unavailable`-only result set is honestly reported as `'fail'` (i.e. "not confirmed to satisfy every selected requirement"), which the report's per-requirement breakdown (`pass`/`fail`/`unavailable` counts) then disambiguates for a human/caller.
|
|
80
|
-
|
|
81
|
-
## Region-property evaluation
|
|
82
|
-
|
|
83
|
-
Reads the bound target's `TargetGeometry` from `targetEvidence`, requires `TargetVisibility.visible === true` (a `bound`-but-hidden target's geometry is treated as unusable, never fabricated as meaningful — see "hidden target" below), converts to both raw and reference-space geometry, and reads `[subject.property]` off each. Exactly Prompt 3's supported vocabulary (`x`/`y`/`width`/`height`/`right`/`bottom`/`centerX`/`centerY`) — no new property was added.
|
|
84
|
-
|
|
85
|
-
## Relationship evaluation
|
|
86
|
-
|
|
87
|
-
Confirms the reference itself exhibits its own selected relationship first (`deriveReferenceRequirementExpectation`'s `matches` field) — if not, `unavailable`/`reference-relationship-not-exhibited` (a reference-authoring problem, never a candidate `fail`). Then resolves both bound targets, calls `deriveLayoutRelationships` once over the whole candidate observation, and looks up the pairwise record for the bound-target pair **scoped to the exact requested relationship family** via an independently-owned, third duplicate of the `RELATIONSHIP_FAMILY_GROUPS` shape (already duplicated by `frontendContractEvaluation.ts` and `externalReferenceRequirements.ts`) — never "the first record for this pair regardless of family," the exact bug class Prompt 3 fixed. A record only derivable in the reversed target order is `unavailable`, never auto-flipped, mirroring Prompt 3's identical reference-side handling. Verified by dedicated tests for pass (M), fail (N), and family-scoping (O, using two regions/targets that simultaneously satisfy multiple relationship families to prove the lookup picks the correct family's record).
|
|
88
|
-
|
|
89
|
-
## Spacing/alignment (measurement) evaluation
|
|
90
|
-
|
|
91
|
-
Supported: exactly Prompt 3's `REFERENCE_REQUIREMENT_MEASUREMENTS` (`vertical-gap`/`horizontal-gap`/`center-x-delta`/`center-y-delta`/`left-edge-delta`/`right-edge-delta`), evaluated via `deriveReferenceRequirementMeasurement` on converted (both raw and reference-space) geometries — no generic geometry expression language was introduced. A geometrically-undefined gap (targets overlapping on the relevant axis) is `unavailable`, mirroring Prompt 3's reference-side treatment of the identical situation. Verified by a dedicated horizontal-gap pass/fail test (behavior P).
|
|
92
|
-
|
|
93
|
-
## Tolerance behavior
|
|
94
|
-
|
|
95
|
-
Reused exactly, never redefined. A single rule, `Math.abs(delta) <= allowed`, covers all three Prompt 3 kinds: `exact` ⇒ `allowed = 0`; `absolute-reference-px` ⇒ `allowed = tolerance.amount` (already in reference-image pixels); `percent` ⇒ `allowed = (tolerance.amount / 100) * Math.abs(referenceValue)`, mirroring v0.5's own `toleranceToPx` "may vary by up to N%" convention exactly, independently reimplemented in reference-image-pixel units (never imported — v0.5's `ContractTolerance` is a different, CSS-pixel-implicit unit). Verified by dedicated tests for `exact`, `absolute-reference-px`, `percent`, and the exact tolerance boundary (behavior C, plus the task's own worked example reproduced verbatim in both the domain and CLI test suites).
|
|
96
|
-
|
|
97
|
-
## Unavailable-evidence behavior
|
|
98
|
-
|
|
99
|
-
Unavailable evidence never becomes a fabricated zero/false/pass/fail anywhere in this module: an unbound/ambiguous/unavailable Prompt 5 binding, an invisible or evidence-less bound target, an unestablishable coordinate scale, and an undervied/not-exhibited reference relationship are each their own explicit `unavailable` result with a distinguishing `reasonCode` — never silently defaulted, never a `fail`.
|
|
100
|
-
|
|
101
|
-
## Compatibility-blocker behavior
|
|
102
|
-
|
|
103
|
-
`evaluateReferenceCandidateCompatibility` is called once; `state === 'incomparable'` short-circuits the whole evaluation to `state: 'not-evaluated'`, `blockedBy: 'incompatible'`, `requirementResults: []`, with the full `ComparabilityResult` embedded verbatim in the `compatibility` field for the caller to inspect the reason. No geometry delta, spacing delta, relationship comparison, or requirement PASS/FAIL is ever computed in this case. Verified by a dedicated test (behavior G, an incompatible candidate viewport).
|
|
104
|
-
|
|
105
|
-
## Binding-blocker behavior
|
|
106
|
-
|
|
107
|
-
`evaluateReferenceRuntimeBindings` is called once (after the compatibility gate has already confirmed a non-incomparable pair, so it is guaranteed not to re-trigger its own internal compatibility short-circuit); its `ok: false` result (a structural declaration/artifact problem) propagates as this function's own `ok: false`. For each requirement, every dependent reference region's binding must be `bound` — `ambiguous`, `unavailable`, or simply absent from the supplied declarations all produce `unavailable`/`binding-unavailable` for that requirement, never a guessed target and never an auto-bind based on geometry or names. Verified by dedicated tests for ambiguous (J), unavailable/not-found (K), and undeclared bindings.
|
|
108
|
-
|
|
109
|
-
## Hidden-target decision
|
|
110
|
-
|
|
111
|
-
A `bound` target (Prompt 5's question: does a stable correspondence exist) whose `TargetVisibility.visible !== true` — including when visibility evidence itself is unavailable — is treated as having no usable geometry: `unavailable`/`candidate-evidence-unavailable`, never a fabricated zero-geometry comparison. This preserves the documented distinction between binding success and fidelity evaluability. Verified by a dedicated test (behavior L).
|
|
112
|
-
|
|
113
|
-
## Persistence decision
|
|
114
|
-
|
|
115
|
-
**No new persisted artifact family.** `evaluateReferenceCandidateFidelity` (and its CLI-facing wrapper) is a pure, on-demand function over already-persisted/in-memory evidence — no `ExternalReferenceFidelityEvaluationArtifact` or equivalent was introduced. Rationale, identical to Prompt 4/5's own precedent: the result is cheap to recompute deterministically from its inputs, and persisting it would invite drift (a re-imported reference, re-observed candidate, or edited binding declaration could silently disagree with a stale persisted record) with no corresponding benefit at this stage. `--enforce`'s exit-code-only effect (mirroring `evaluate-contract` exactly) was judged sufficient CLI-side signal without persistence. This may be revisited only if Prompt 7's architecture (bounded visual-fidelity mismatch projection, v0.6 bounded-agent-context integration) proves persistence necessary — not assumed here, and no `BLOCKED_ESCALATE_TO_FULL_STAGE_CONTEXT` condition was triggered by this decision since the existing CLI could return a fully structured result without it.
|
|
116
|
-
|
|
117
|
-
## Programmatic interface
|
|
118
|
-
|
|
119
|
-
Exported from `src/index.ts` (additive): types `ReferenceRequirementFidelityStatus`, `ReferenceRequirementFidelityReasonCode`, `ReferenceRequirementFidelityResult`, `ReferenceFidelityState`, `ReferenceFidelityBlockReason`, `ReferenceCandidateFidelityEvaluation`, `EvaluateReferenceCandidateFidelityResult`, `EvaluateReferenceCandidateFidelityOptions`, `EvaluateReferenceFidelityOptions`, `ApplicationReferenceFidelityResult`; values `REFERENCE_REQUIREMENT_FIDELITY_STATUSES`, `REFERENCE_REQUIREMENT_FIDELITY_REASON_CODES`, `REFERENCE_FIDELITY_STATES`, `REFERENCE_FIDELITY_BLOCK_REASONS`, `evaluateReferenceCandidateFidelity`, `evaluateReferenceCandidateFidelityFromArtifactRoots`. No internal helper (`evaluateOneRequirement`, `deriveCoordinateScale`, `resolveConfiguredTargetName`, etc.) is exported.
|
|
120
|
-
|
|
121
|
-
## CLI surface
|
|
122
|
-
|
|
123
|
-
New command: `evaluate-reference-fidelity --reference <root> --candidate <root> [--bindings-file <json-file>] [--enforce]`. `--bindings-file` follows the exact `--requirements-file`/`--regions-file` wrapped-object convention (`{ "bindings": [...] }`, root-field allowlist enforced, unwrapped content handed straight to the domain validator). CLI code (`parseEvaluateReferenceFidelityArgs`, `loadBindingsFile`, `runEvaluateReferenceFidelityCommand`) owns only flag syntax, duplicate/missing-flag detection, file reading, JSON parsing, root-shape validation, output presentation, and exit-code selection — every semantic rule (adequacy, compatibility, binding, coordinate mapping, tolerance, relationship) stays owned by the domain module, and `evaluateReferenceCandidateFidelityFromArtifactRoots` owns artifact-root reading/orchestration. `TOP_LEVEL_HELP` and `runCli`'s dispatch table were updated; `cli.ts`'s existing "no browser/persistence implementation of its own" regression test (`B5-TST-008,009`) continues to pass unchanged, confirming the new command imports neither `artifacts/` nor a filesystem-write function.
|
|
124
|
-
|
|
125
|
-
## CLI exit behavior
|
|
126
|
-
|
|
127
|
-
Mirrors `evaluate-contract`'s exact precedent: `--enforce` is applied only after the fidelity evaluation has already been computed, and affects only the process exit status for an already-final `state: 'fail'` result (never its printed content). `state: 'not-evaluated'` always exits `0` regardless of `--enforce` — a reference-adequacy or compatibility blocker is a successful, honest non-evaluation, never treated as a design mismatch or an execution error. An actual operational failure (invalid syntax, unreadable/malformed `--reference`/`--candidate`, malformed `--bindings-file`, or an invalid binding declaration) always exits nonzero regardless of `--enforce`. Verified end-to-end by three CLI tests reproducing pass/fail/not-evaluated with and without `--enforce`.
|
|
128
|
-
|
|
129
|
-
## Boundedness
|
|
130
|
-
|
|
131
|
-
Reuses existing bounds without introducing new ones: `reference.requirements` is already bounded to ≤50 (`MAX_REFERENCE_REQUIREMENTS`, Prompt 3) at construction time, so `requirementResults` is bounded identically (1:1 with requirements); `bindingDeclarations` is already bounded to ≤20 (`MAX_REFERENCE_RUNTIME_BINDINGS`, Prompt 5); `deriveLayoutRelationships`'s own `MAX_CONFIGURED_TARGETS_FOR_RELATIONSHIPS`/`MAX_PAIRWISE_RELATIONSHIP_RECORDS` bounds (v0.4) apply unchanged, since it is called exactly once per relationship requirement over the same candidate. No unbounded diff/diagnostic collection is ever produced — Prompt 6 evaluates only explicitly authored Prompt 3 requirements, never "all visible differences."
|
|
132
|
-
|
|
133
|
-
## Provenance
|
|
134
|
-
|
|
135
|
-
Each `ReferenceRequirementFidelityResult` carries `requirementId`, `category`, `expectedDependentMode` (when present), `subject` (which itself names the reference region(s)/property/relationship/measurement), `boundRuntimeTargets`, and either the numeric (`referenceValue`/`candidateRawValue`/`candidateValue`/`delta`/`tolerance`) or relationship (`expectedRelationship`/`actualRelationship`) evidence that produced its status. The top-level `ReferenceCandidateFidelityEvaluation` additionally carries `referenceId`/`referenceRequestId`/`candidateObservationId`/`candidateRequestId`, the full Prompt 3 `adequacy`, the full Prompt 4 `compatibility` (when computed), and the full Prompt 5 `bindings` evaluation (when computed) — nothing is an unexplained delta.
|
|
136
|
-
|
|
137
|
-
## Files changed
|
|
138
|
-
|
|
139
|
-
New:
|
|
140
|
-
- `src/domain/externalReferenceFidelity.ts`
|
|
141
|
-
- `src/application/referenceFidelityEvaluationService.ts`
|
|
142
|
-
- `tests/unit/externalReferenceFidelity.test.ts`
|
|
143
|
-
- `tests/unit/cliEvaluateReferenceFidelity.test.ts`
|
|
144
|
-
|
|
145
|
-
Modified:
|
|
146
|
-
- `src/cli.ts` (`evaluate-reference-fidelity` command: help text, argument parser, `loadBindingsFile`, run function, `TOP_LEVEL_HELP`, dispatch table)
|
|
147
|
-
- `src/index.ts` (public export surface for the above)
|
|
148
|
-
- `docs/ARCHITECTURE.md`, `docs/CONTRACTS.md`, `docs/WORKFLOWS.md`, `docs/COMMANDS.md`
|
|
149
|
-
|
|
150
|
-
## Tests
|
|
151
|
-
|
|
152
|
-
927 unit tests total (889 pre-existing + 38 new), all pure domain/application/CLI-level tests — zero Chromium launches for fidelity evaluation itself (the one CLI end-to-end test constructs its candidate `ObservationArtifact` by writing a `manifest.json` directly to disk, exactly as `tests/unit/artifactWriter.test.ts` already does for writer-focused tests, never via `observe()`/a real browser).
|
|
153
|
-
|
|
154
|
-
- `tests/unit/externalReferenceFidelity.test.ts` (29 tests): behaviors A–W from the task's own behavior model, including the exact worked example (424 reference px, 223→446 candidate px, delta 22, tolerance ±4, FAIL), the exact tolerance boundary (428 passes, 428.0001 fails), `exact`/`absolute-reference-px`/`percent` tolerance kinds, 2x coordinate scale, aspect-ratio-mismatch and missing-viewport coordinate-mapping unavailability, compatibility/adequacy blockers, bound/ambiguous/unavailable/undeclared bindings, hidden-target unavailability, relationship pass/fail/family-scoping and not-exhibited-by-reference, a supported numeric measurement (horizontal-gap), deterministic ordering and pure-function repeatability, all-pass/one-fail/one-unavailable overall results, zero-requirements inadequacy, full three-input immutability, fail-closed structural validation, absence of source-ownership fields, and a source-scan confirming no browser/filesystem import.
|
|
155
|
-
- `tests/unit/cliEvaluateReferenceFidelity.test.ts` (9 tests): help text, top-level help listing, required-flag errors, missing `--reference`/`--candidate` errors, malformed `--bindings-file` (bad JSON/wrong root shape/unknown field), and three full end-to-end runs (pass, fail, not-evaluated) each checked with and without `--enforce`, reproducing the task's exact worked example through the real CLI/application/domain stack.
|
|
156
|
-
|
|
157
|
-
## Validation results
|
|
158
|
-
|
|
159
|
-
All commands run from the repository root, after the implementation commit:
|
|
160
|
-
|
|
161
|
-
- `npm run typecheck` — pass, zero errors.
|
|
162
|
-
- `npm run lint` — pass, zero errors/warnings.
|
|
163
|
-
- `npm test` — 48 test files, 927 tests, all pass.
|
|
164
|
-
- `npm run build` — pass, clean `tsc` compile.
|
|
165
|
-
- `npm run check:docs` — pass (17 required files present, `ROADMAP.md` format intact — no implementation batches added).
|
|
166
|
-
- `git diff --check` — exit 0, no whitespace errors.
|
|
167
|
-
- `npm pack --dry-run` — pass; `dist/domain/externalReferenceFidelity.{js,d.ts,js.map}` and `dist/application/referenceFidelityEvaluationService.{js,d.ts,js.map}` confirmed present in the tarball listing.
|
|
168
|
-
- `npm run test:security` — pass (5 + 63 = 68 tests), run because a new CLI/config-file input surface (`--bindings-file`) was added.
|
|
169
|
-
- `npm run test:browser` — pass on re-run (9 files, 120 tests). One transient failure (`v0.2: an ambiguous locator stops immediately...`) occurred on the first full-suite run; re-running that single test in isolation passed immediately, and a full second `test:browser` run passed 120/120 — confirmed environmental flakiness in a real-Chromium timing-sensitive test unrelated to any file this prompt touched (Prompt 6 never imports or modifies `chromiumAdapter.ts`/`evidenceCapture.ts`), not a regression.
|
|
170
|
-
|
|
171
|
-
## Regression results
|
|
172
|
-
|
|
173
|
-
- Prompt 1 external-reference foundation: unaffected — `externalReference.ts` untouched (only its already-exported type guards were imported).
|
|
174
|
-
- Prompt 2 regions/shared relationships: unaffected — `externalReferenceRegions.ts`/`externalReferenceRegionRelationships.ts` untouched.
|
|
175
|
-
- Prompt 3 requirements/tolerances/adequacy: unaffected — `externalReferenceRequirements.ts` untouched; its exported functions were called, never modified.
|
|
176
|
-
- Prompt 4 compatibility: unaffected — `externalReferenceCompatibility.ts`/`explicitState.ts`/`externalReferenceApplicability.ts` untouched.
|
|
177
|
-
- Prompt 5 binding: unaffected — `externalReferenceRuntimeBinding.ts` untouched.
|
|
178
|
-
- v0.4 comparison/relationship behavior: unaffected — `comparisonEngine.ts`/`relationships.ts` untouched (only already-exported functions were called).
|
|
179
|
-
- v0.5 contracts/evaluation: unaffected — no file in `frontendContracts*.ts` touched.
|
|
180
|
-
- v0.6 bounded context/correlation: unaffected — `boundedAgentContext*.ts` untouched; read for precedent only in an earlier prompt, not this one.
|
|
181
|
-
- Full unit (927/927), browser (120/120 on the confirming re-run), and security (68/68) suites all pass with zero regressions.
|
|
182
|
-
|
|
183
|
-
## Security impact
|
|
184
|
-
|
|
185
|
-
- No new external input surface beyond a local, caller-controlled JSON file (`--bindings-file`) — the same trust boundary as every existing `--*-file` flag.
|
|
186
|
-
- `evaluateReferenceCandidateFidelity` performs no filesystem access, no network access, and no browser/Chromium invocation of any kind; `evaluateReferenceCandidateFidelityFromArtifactRoots` performs only read-only artifact-manifest reads through the existing readers.
|
|
187
|
-
- `test:security` (policy + real-Chromium adapter tests) re-run and passing, confirming no regression to the existing safety/navigation policy surface (untouched by this prompt).
|
|
188
|
-
|
|
189
|
-
## Documentation changes
|
|
190
|
-
|
|
191
|
-
- `docs/CONTRACTS.md` — new "v0.7 Prompt 6 structured reference-vs-candidate fidelity evaluation" section (full type shapes, evaluation order, coordinate-mapping/aspect-ratio rules, tolerance reuse, region-property/measurement/relationship evaluation, binding gate, category preservation, overall-result rule, persistence decision, CLI summary).
|
|
192
|
-
- `docs/ARCHITECTURE.md` — new paragraph in the "Planned v0.7–v0.10" section describing the Prompt 6 module additions and their reuse of Prompts 3/4/5 and v0.4.
|
|
193
|
-
- `docs/WORKFLOWS.md` — "Current external-reference foundation workflow" retitled to "Prompts 1-6" and extended with the `evaluate-reference-fidelity` command flow and test-coverage summary.
|
|
194
|
-
- `docs/COMMANDS.md` — new `## `evaluate-reference-fidelity`` section, mirroring the existing `## `evaluate-contract`` section's structure exactly.
|
|
195
|
-
|
|
196
|
-
## Tooling incidents
|
|
197
|
-
|
|
198
|
-
None. No orchestrator was invoked (direct-implementation mode used throughout, consistent with Prompts 2–6); no background/speculative subagent writes occurred; the Prompt 1 stray-fork-writes stash remains untouched, unapplied, and unmined as precedent. The one transient real-Chromium test failure encountered during validation is documented under Validation/Regression results above as confirmed flakiness, not a tooling or implementation defect.
|
|
199
|
-
|
|
200
|
-
## Out-of-scope confirmation
|
|
201
|
-
|
|
202
|
-
This prompt implements no bounded coding-agent correction packet, no v0.6 bounded-agent-context change, no source/static correlation change, no coding-agent invocation, no source editing, no rerender/retry loop, no baseline/per-change overall workflow composition, no viewer, no annotation, no automatic target matching, no pixel/image similarity (no SSIM/perceptual hash/OCR/computer vision of any kind), and no new style-fidelity family (background color, text color, font size/weight, line height, border radius, shadow, gradient, opacity, icons) beyond Prompt 3's already-frozen requirement vocabulary. `evaluateReferenceCandidateFidelity` never attaches `sourceOwner`/`sourceFile`/`component`/`symbol`/`causedBy` to any result (verified by a dedicated test scanning the serialized evaluation for those exact terms).
|
|
203
|
-
|
|
204
|
-
## Known limitations
|
|
205
|
-
|
|
206
|
-
- The coordinate-mapping model supports only a single, deliberately bounded full-frame scale per evaluation — no per-region cropping, offset, rotation, or perspective mapping exists or was attempted, per the task's own explicit prohibition.
|
|
207
|
-
- `expectedDependentMode` (`required`/`permitted`) is carried through as provenance only and never changes the pass/fail rule, since Prompt 3 itself never implemented a directional evaluation difference for its own expectation/adequacy derivation (unlike v0.5's runtime-directional contract clauses) — documented explicitly rather than silently ignored.
|
|
208
|
-
- No `--output`/persistence flag exists for `evaluate-reference-fidelity`; a caller needing to save a fidelity result must do so itself (e.g. redirecting stdout, or consuming the programmatic API directly) until/unless a later prompt's architecture proves in-repository persistence necessary.
|
|
209
|
-
|
|
210
|
-
## Remaining risks
|
|
211
|
-
|
|
212
|
-
- None identified that block this prompt's own scope. The primary forward consideration for Prompt 7 (bounded visual-fidelity mismatch projection and v0.6 bounded-agent-context integration) is how it will attach `ReferenceRequirementFidelityResult`/`ReferenceCandidateFidelityEvaluation` evidence to `RuntimeStaticCorrelationRecord`-shaped context without conflating the two identity domains this prompt was careful to keep separate (reference-region/runtime-target vs. runtime-target/static-source) — flagged for that prompt's own precedent review, not preempted here.
|
|
213
|
-
|
|
214
|
-
## Exact next action
|
|
215
|
-
|
|
216
|
-
v0.7 Prompt 7 — bounded visual-fidelity mismatch projection and v0.6 bounded-agent-context integration.
|
|
1
|
+
# v0.7 Prompt 6 — Structured Reference-vs-Candidate Fidelity Evaluation
|
|
2
|
+
|
|
3
|
+
**VERDICT: PASS_V0_7_REFERENCE_FIDELITY_PROMPT6**
|
|
4
|
+
|
|
5
|
+
## Repository / branch / heads
|
|
6
|
+
|
|
7
|
+
- Repository: `my-frontend-observer` (path: `Z:\Users\newuser\Projects\my-frontend-observer`)
|
|
8
|
+
- Branch: `implementation/v0.7-reference-fidelity`, branched from the exact completed Prompt 5 HEAD
|
|
9
|
+
- Starting HEAD (branch point, Prompt 5 report commit): `74c710daa62cd5a56f1b1b2eac3a43d24dfe69fb`
|
|
10
|
+
- Prompt 5 base HEAD confirmed to contain both required commits: `74cc148` (implementation) and `74c710d` (report), and the full completed Prompt 1–4 lineage — verified via `git log --oneline` before branching.
|
|
11
|
+
- Implementation commit (this prompt): `91f5e736a1a8472984b0c8c7a7e179e307f1a5ce` — "Add v0.7 Prompt 6 structured reference-vs-candidate fidelity evaluation"
|
|
12
|
+
- Ending HEAD: the report commit that follows this file's commit.
|
|
13
|
+
|
|
14
|
+
## Git status
|
|
15
|
+
|
|
16
|
+
Preflight (`git status --short`) showed a clean working tree at the exact Prompt 5 report HEAD; `git stash list` showed exactly the one preserved Prompt 1 stray-fork-writes entry. The branch was created with `git checkout -b implementation/v0.7-reference-fidelity` and ancestry verified with `git merge-base --is-ancestor 74c710d HEAD`. No reset/clean/discard operation was used; the Prompt 1 stash was neither applied nor dropped. Post-implementation `git status --short` is clean except for this report file (staged and committed separately, per convention).
|
|
17
|
+
|
|
18
|
+
## Tooling
|
|
19
|
+
|
|
20
|
+
- Resolved latest published `@dailephd/my-dev-kit`: `1.12.3`, pinned exactly via `npx -y @dailephd/my-dev-kit@1.12.3`.
|
|
21
|
+
- Fresh Prompt 6 repository index built under a repository-local, gitignored root: `.my-dev-kit-context/index-prompt6/` (57 files, 868 symbols indexed) — not reused from Prompt 1–5's indexes.
|
|
22
|
+
|
|
23
|
+
## Prompt 1–5 contracts verified
|
|
24
|
+
|
|
25
|
+
Read directly (source, not memory) before writing any code:
|
|
26
|
+
|
|
27
|
+
- `src/domain/externalReferenceRequirements.ts` — the full Prompt 3 requirement/tolerance/expectation/adequacy model, including the exact `RELATIONSHIP_FAMILY_GROUPS`-scoped lookup fix and `deriveReferenceRequirementAdequacy`'s "100% unavailable requirements ⇒ inadequate" rule.
|
|
28
|
+
- `src/domain/externalReferenceRegions.ts` — `ReferenceRegionGeometry`'s exact field set (`x`/`y`/`width`/`height`/`right`/`bottom`/`centerX`/`centerY`), reused as the shared shape for both reference and (converted) candidate geometry.
|
|
29
|
+
- `src/domain/externalReference.ts` — `isImportedExternalReferenceArtifact`/`isApprovedExternalReferenceArtifact` type guards, reused to read image dimensions uniformly across both lifecycle variants without duplicating validation logic.
|
|
30
|
+
- `src/domain/externalReferenceCompatibility.ts` (Prompt 4) — `evaluateReferenceCandidateCompatibility`'s exact signature and result shape, reused verbatim as the compatibility gate.
|
|
31
|
+
- `src/domain/externalReferenceRuntimeBinding.ts` (Prompt 5) — `evaluateReferenceRuntimeBindings`'s exact signature/result shape and its own case-insensitive target-name-resolution convention (mirrored, not imported, per the established per-module small-helper convention).
|
|
32
|
+
- `src/domain/schema.ts` (v0.2) — `TargetGeometry`, `TargetVisibility`, `TargetResolution`/`TargetSelectionStatus`, and `ObservationArtifact.targetEvidence`'s exact-configured-name keying.
|
|
33
|
+
- `src/domain/comparisonEngine.ts` (v0.4) — `targetPresence` (already additively exported by Prompt 5), reused rather than re-derived.
|
|
34
|
+
- `src/domain/relationships.ts` (v0.4) — `deriveLayoutRelationships`, `PairwiseLayoutRelationship`'s `subjectTarget`/`relatedTarget` fields, `UnresolvedRelationshipTarget`/`UNRESOLVED_TARGET_REASONS` (confirming hidden targets are already excluded from relationship derivation by v0.4 itself), and every pairwise relationship-family constant.
|
|
35
|
+
- `src/domain/frontendContracts.ts`/`frontendContractEvaluation.ts` (v0.5) — `CLAUSE_RESULT_STATUSES`/`ClauseEvaluationResult` (the `pass`/`fail`/`unavailable`/`conflict` precedent), `evaluateFrontendContract`'s overall-verdict aggregation rule (`every clause pass && no unexpected changes ⇒ PASS`), `clauseMode`'s `required`/`permitted` handling (confirming it is a v0.5-only, directional-primitive-specific concept never mirrored in Prompt 3's own expectation/adequacy derivation), and `toleranceToPx`'s percent-denominator convention (`(amount/100) * abs(beforeValue)`), reused independently for reference-image-pixel units.
|
|
36
|
+
- `src/application/comparisonService.ts` — the exact `compareAndPersistFromArtifactRoots` application-layer shape (read-both-artifacts-then-delegate-to-the-pure-function), mirrored for `evaluateReferenceCandidateFidelityFromArtifactRoots`.
|
|
37
|
+
- `src/cli.ts` — the full `evaluate-contract` command (help text, argument parser, `--enforce` exit-code handling) as the direct precedent for `evaluate-reference-fidelity`, and every existing `--*-file` loader (`loadRegionsFile`/`loadRequirementsFile`/`loadApplicabilityFile`) as the precedent for `loadBindingsFile`.
|
|
38
|
+
|
|
39
|
+
No architecture blocker was hit; no `BLOCKED_ESCALATE_TO_FULL_STAGE_CONTEXT` condition applied at any point.
|
|
40
|
+
|
|
41
|
+
## Precedent review outcomes
|
|
42
|
+
|
|
43
|
+
- **A. Result vocabulary** — `pass`/`fail`/`unavailable` reused as the same honest three-state shape v0.5's `CLAUSE_RESULT_STATUSES` already established, but as an independently-owned constant (`REFERENCE_REQUIREMENT_FIDELITY_STATUSES`) deliberately excluding `'conflict'` — Prompt 6 has no cross-requirement authoring-conflict concept (each requirement is evaluated independently against its own subject).
|
|
44
|
+
- **B. Overall fidelity** — a compatible/evaluable pair reaches `state: 'pass'` only when every requirement result is `'pass'`; any `'fail'`/`'unavailable'` forces `'fail'`. An incompatible or reference-inadequate pair never reaches that computation at all — a distinct `not-evaluated` state, with `blockedBy: 'reference-inadequate' | 'incompatible'`, is returned instead.
|
|
45
|
+
- **C. Relationship evaluation** — reuses Prompt 2's `deriveReferenceRequirementExpectation` (reference side) and v0.4's `deriveLayoutRelationships` (candidate side) exactly; no duplicated geometry predicate exists anywhere in the new module.
|
|
46
|
+
- **D. Value derivation** — reference expected values always come from `deriveReferenceRequirementExpectation`/`deriveReferenceRequirementMeasurement` (Prompt 3), never recomputed by hand; candidate values always come from `ObservationArtifact.targetEvidence`'s existing `TargetGeometry`/`TargetVisibility` evidence, never a second derivation.
|
|
47
|
+
- **E. Persistence** — no new persisted artifact family; documented decision below.
|
|
48
|
+
|
|
49
|
+
## Fidelity evaluator type/function names
|
|
50
|
+
|
|
51
|
+
- `domain/externalReferenceFidelity.ts`: `evaluateReferenceCandidateFidelity(reference, candidate, bindingDeclarations, options?)` → `{ ok: true; evaluation: ReferenceCandidateFidelityEvaluation } | { ok: false; reason: string }`.
|
|
52
|
+
- `application/referenceFidelityEvaluationService.ts`: `evaluateReferenceCandidateFidelityFromArtifactRoots(referenceRoot, candidateRoot, bindingDeclarations, options?)` — the CLI-facing, artifact-root-reading counterpart.
|
|
53
|
+
- CLI: `evaluate-reference-fidelity`.
|
|
54
|
+
|
|
55
|
+
## Evaluation ordering / gates
|
|
56
|
+
|
|
57
|
+
Frozen order, exactly as specified and never reordered: reference structural validation (`isValidExternalReferenceArtifact`) → candidate structural validation (`isValidObservationArtifact`) → binding-declaration structural validation (`isValidReferenceRuntimeBindingDeclarations`, invoked inside the Prompt 5 evaluator) → Prompt 3 reference adequacy (`deriveReferenceRequirementAdequacy`) → Prompt 4 compatibility (`evaluateReferenceCandidateCompatibility`) → Prompt 5 binding evaluation (`evaluateReferenceRuntimeBindings`) → per-requirement candidate-evidence/coordinate-mapping checks → per-requirement tolerance/relationship comparison → overall result.
|
|
58
|
+
|
|
59
|
+
The first three (structural) failures return `{ ok: false, reason }` — a caller/config error, never a fidelity outcome. `adequacy.status === 'inadequate'` and `compatibility.state === 'incomparable'` each short-circuit to `state: 'not-evaluated'` with `requirementResults: []` — verified by dedicated tests (behaviors G/H). `adequacy.status === 'partial'` does **not** block evaluation; it proceeds normally, and the specific reference-side-unavailable requirements independently re-derive their own `unavailable` result at the per-requirement stage (not looked up from the adequacy result — verified by the "reference itself does not exhibit" test, which requires a second, evaluable requirement to keep adequacy at `partial` rather than `inadequate`, confirming the two computations are genuinely independent).
|
|
60
|
+
|
|
61
|
+
## Coordinate-mapping model
|
|
62
|
+
|
|
63
|
+
One explicit, deterministic full-frame scale: `scaleX = referenceImageWidth / applicableViewportWidth`, `scaleY = referenceImageHeight / applicableViewportHeight`, derived once per evaluation call from `reference.applicability.viewport` (Prompt 4) and the reference image's own pixel dimensions (Prompt 1, read uniformly across both lifecycle variants via the existing `isImportedExternalReferenceArtifact`/`isApprovedExternalReferenceArtifact` type guards). Candidate `TargetGeometry` (CSS pixels) is reshaped into a `ReferenceRegionGeometry`-shaped value (adding `centerX`/`centerY`, computed identically to `deriveReferenceRegionGeometry`) both raw and scaled; horizontal fields (`x`/`width`/`right`/`centerX`) always scale by `scaleX`, vertical fields (`y`/`height`/`bottom`/`centerY`) always scale by `scaleY` — this falls out automatically from the per-field conversion, never a hand-picked per-property axis table. `region-measurement` subjects reuse Prompt 3's exact `deriveReferenceRequirementMeasurement` on the converted geometries directly, rather than a parallel "runtime version" of the same gap/delta formulas.
|
|
64
|
+
|
|
65
|
+
## Image/viewport scale rule and aspect-ratio rule
|
|
66
|
+
|
|
67
|
+
`ASPECT_RATIO_MAPPING_TOLERANCE = 0.01` (1% relative difference between `scaleX` and `scaleY`) is an independently-owned, tiny coordinate-mapping-validity constant — never a user-authored Prompt 3 design tolerance. If `scaleX`/`scaleY` disagree beyond this bound, or `reference.applicability.viewport` is absent entirely, `deriveCoordinateScale` returns `{ ok: false, reason }` and every numeric (`region-property`/`region-measurement`) requirement in that evaluation becomes `unavailable`/`coordinate-mapping-unavailable` — verified by dedicated tests (behaviors E and F). No cropping, offset, rotation, or perspective registration is implemented anywhere; only this single bounded full-frame check. Categorical `region-relationship` requirements are unaffected by an unavailable scale (they never call `deriveCoordinateScale` at all).
|
|
68
|
+
|
|
69
|
+
## Numeric precision rule
|
|
70
|
+
|
|
71
|
+
No intermediate rounding anywhere in the arithmetic path (`referenceValue`, `candidateRawValue`, `candidateValue`, `delta` are all plain JS `number`s carried through at full floating-point precision). Tolerance comparison is a single `Math.abs(delta) <= allowed` check with no added epsilon — verified by the exact-boundary test (a delta equal to the tolerance passes; a delta a tiny fraction over it fails), matching the task's own worked boundary example (428 passes, 428.0001 fails against a reference of 424 ± 4).
|
|
72
|
+
|
|
73
|
+
## Requirement result vocabulary
|
|
74
|
+
|
|
75
|
+
`REFERENCE_REQUIREMENT_FIDELITY_STATUSES = ['pass', 'fail', 'unavailable']`. `reasonCode` (one of `reference-evidence-unavailable`, `reference-relationship-not-exhibited`, `binding-unavailable`, `candidate-evidence-unavailable`, `coordinate-mapping-unavailable`) and `detail` are present only when `status === 'unavailable'` — a `fail` is already fully explained by its numeric (`referenceValue`/`candidateRawValue`/`candidateValue`/`delta`/`tolerance`) or relationship (`expectedRelationship`/`actualRelationship`) fields, so no reason code is needed there.
|
|
76
|
+
|
|
77
|
+
## Overall fidelity vocabulary
|
|
78
|
+
|
|
79
|
+
`REFERENCE_FIDELITY_STATES = ['not-evaluated', 'pass', 'fail']`, with `REFERENCE_FIDELITY_BLOCK_REASONS = ['reference-inadequate', 'incompatible']` populated only when `state === 'not-evaluated'`. There is no separate "partial pass" bucket once evaluation has actually run: any `unavailable` requirement result forces `state: 'fail'`, satisfying "any required fail or unavailable must prevent overall PASS" (section 34) without inventing a fourth top-level state — an `unavailable`-only result set is honestly reported as `'fail'` (i.e. "not confirmed to satisfy every selected requirement"), which the report's per-requirement breakdown (`pass`/`fail`/`unavailable` counts) then disambiguates for a human/caller.
|
|
80
|
+
|
|
81
|
+
## Region-property evaluation
|
|
82
|
+
|
|
83
|
+
Reads the bound target's `TargetGeometry` from `targetEvidence`, requires `TargetVisibility.visible === true` (a `bound`-but-hidden target's geometry is treated as unusable, never fabricated as meaningful — see "hidden target" below), converts to both raw and reference-space geometry, and reads `[subject.property]` off each. Exactly Prompt 3's supported vocabulary (`x`/`y`/`width`/`height`/`right`/`bottom`/`centerX`/`centerY`) — no new property was added.
|
|
84
|
+
|
|
85
|
+
## Relationship evaluation
|
|
86
|
+
|
|
87
|
+
Confirms the reference itself exhibits its own selected relationship first (`deriveReferenceRequirementExpectation`'s `matches` field) — if not, `unavailable`/`reference-relationship-not-exhibited` (a reference-authoring problem, never a candidate `fail`). Then resolves both bound targets, calls `deriveLayoutRelationships` once over the whole candidate observation, and looks up the pairwise record for the bound-target pair **scoped to the exact requested relationship family** via an independently-owned, third duplicate of the `RELATIONSHIP_FAMILY_GROUPS` shape (already duplicated by `frontendContractEvaluation.ts` and `externalReferenceRequirements.ts`) — never "the first record for this pair regardless of family," the exact bug class Prompt 3 fixed. A record only derivable in the reversed target order is `unavailable`, never auto-flipped, mirroring Prompt 3's identical reference-side handling. Verified by dedicated tests for pass (M), fail (N), and family-scoping (O, using two regions/targets that simultaneously satisfy multiple relationship families to prove the lookup picks the correct family's record).
|
|
88
|
+
|
|
89
|
+
## Spacing/alignment (measurement) evaluation
|
|
90
|
+
|
|
91
|
+
Supported: exactly Prompt 3's `REFERENCE_REQUIREMENT_MEASUREMENTS` (`vertical-gap`/`horizontal-gap`/`center-x-delta`/`center-y-delta`/`left-edge-delta`/`right-edge-delta`), evaluated via `deriveReferenceRequirementMeasurement` on converted (both raw and reference-space) geometries — no generic geometry expression language was introduced. A geometrically-undefined gap (targets overlapping on the relevant axis) is `unavailable`, mirroring Prompt 3's reference-side treatment of the identical situation. Verified by a dedicated horizontal-gap pass/fail test (behavior P).
|
|
92
|
+
|
|
93
|
+
## Tolerance behavior
|
|
94
|
+
|
|
95
|
+
Reused exactly, never redefined. A single rule, `Math.abs(delta) <= allowed`, covers all three Prompt 3 kinds: `exact` ⇒ `allowed = 0`; `absolute-reference-px` ⇒ `allowed = tolerance.amount` (already in reference-image pixels); `percent` ⇒ `allowed = (tolerance.amount / 100) * Math.abs(referenceValue)`, mirroring v0.5's own `toleranceToPx` "may vary by up to N%" convention exactly, independently reimplemented in reference-image-pixel units (never imported — v0.5's `ContractTolerance` is a different, CSS-pixel-implicit unit). Verified by dedicated tests for `exact`, `absolute-reference-px`, `percent`, and the exact tolerance boundary (behavior C, plus the task's own worked example reproduced verbatim in both the domain and CLI test suites).
|
|
96
|
+
|
|
97
|
+
## Unavailable-evidence behavior
|
|
98
|
+
|
|
99
|
+
Unavailable evidence never becomes a fabricated zero/false/pass/fail anywhere in this module: an unbound/ambiguous/unavailable Prompt 5 binding, an invisible or evidence-less bound target, an unestablishable coordinate scale, and an undervied/not-exhibited reference relationship are each their own explicit `unavailable` result with a distinguishing `reasonCode` — never silently defaulted, never a `fail`.
|
|
100
|
+
|
|
101
|
+
## Compatibility-blocker behavior
|
|
102
|
+
|
|
103
|
+
`evaluateReferenceCandidateCompatibility` is called once; `state === 'incomparable'` short-circuits the whole evaluation to `state: 'not-evaluated'`, `blockedBy: 'incompatible'`, `requirementResults: []`, with the full `ComparabilityResult` embedded verbatim in the `compatibility` field for the caller to inspect the reason. No geometry delta, spacing delta, relationship comparison, or requirement PASS/FAIL is ever computed in this case. Verified by a dedicated test (behavior G, an incompatible candidate viewport).
|
|
104
|
+
|
|
105
|
+
## Binding-blocker behavior
|
|
106
|
+
|
|
107
|
+
`evaluateReferenceRuntimeBindings` is called once (after the compatibility gate has already confirmed a non-incomparable pair, so it is guaranteed not to re-trigger its own internal compatibility short-circuit); its `ok: false` result (a structural declaration/artifact problem) propagates as this function's own `ok: false`. For each requirement, every dependent reference region's binding must be `bound` — `ambiguous`, `unavailable`, or simply absent from the supplied declarations all produce `unavailable`/`binding-unavailable` for that requirement, never a guessed target and never an auto-bind based on geometry or names. Verified by dedicated tests for ambiguous (J), unavailable/not-found (K), and undeclared bindings.
|
|
108
|
+
|
|
109
|
+
## Hidden-target decision
|
|
110
|
+
|
|
111
|
+
A `bound` target (Prompt 5's question: does a stable correspondence exist) whose `TargetVisibility.visible !== true` — including when visibility evidence itself is unavailable — is treated as having no usable geometry: `unavailable`/`candidate-evidence-unavailable`, never a fabricated zero-geometry comparison. This preserves the documented distinction between binding success and fidelity evaluability. Verified by a dedicated test (behavior L).
|
|
112
|
+
|
|
113
|
+
## Persistence decision
|
|
114
|
+
|
|
115
|
+
**No new persisted artifact family.** `evaluateReferenceCandidateFidelity` (and its CLI-facing wrapper) is a pure, on-demand function over already-persisted/in-memory evidence — no `ExternalReferenceFidelityEvaluationArtifact` or equivalent was introduced. Rationale, identical to Prompt 4/5's own precedent: the result is cheap to recompute deterministically from its inputs, and persisting it would invite drift (a re-imported reference, re-observed candidate, or edited binding declaration could silently disagree with a stale persisted record) with no corresponding benefit at this stage. `--enforce`'s exit-code-only effect (mirroring `evaluate-contract` exactly) was judged sufficient CLI-side signal without persistence. This may be revisited only if Prompt 7's architecture (bounded visual-fidelity mismatch projection, v0.6 bounded-agent-context integration) proves persistence necessary — not assumed here, and no `BLOCKED_ESCALATE_TO_FULL_STAGE_CONTEXT` condition was triggered by this decision since the existing CLI could return a fully structured result without it.
|
|
116
|
+
|
|
117
|
+
## Programmatic interface
|
|
118
|
+
|
|
119
|
+
Exported from `src/index.ts` (additive): types `ReferenceRequirementFidelityStatus`, `ReferenceRequirementFidelityReasonCode`, `ReferenceRequirementFidelityResult`, `ReferenceFidelityState`, `ReferenceFidelityBlockReason`, `ReferenceCandidateFidelityEvaluation`, `EvaluateReferenceCandidateFidelityResult`, `EvaluateReferenceCandidateFidelityOptions`, `EvaluateReferenceFidelityOptions`, `ApplicationReferenceFidelityResult`; values `REFERENCE_REQUIREMENT_FIDELITY_STATUSES`, `REFERENCE_REQUIREMENT_FIDELITY_REASON_CODES`, `REFERENCE_FIDELITY_STATES`, `REFERENCE_FIDELITY_BLOCK_REASONS`, `evaluateReferenceCandidateFidelity`, `evaluateReferenceCandidateFidelityFromArtifactRoots`. No internal helper (`evaluateOneRequirement`, `deriveCoordinateScale`, `resolveConfiguredTargetName`, etc.) is exported.
|
|
120
|
+
|
|
121
|
+
## CLI surface
|
|
122
|
+
|
|
123
|
+
New command: `evaluate-reference-fidelity --reference <root> --candidate <root> [--bindings-file <json-file>] [--enforce]`. `--bindings-file` follows the exact `--requirements-file`/`--regions-file` wrapped-object convention (`{ "bindings": [...] }`, root-field allowlist enforced, unwrapped content handed straight to the domain validator). CLI code (`parseEvaluateReferenceFidelityArgs`, `loadBindingsFile`, `runEvaluateReferenceFidelityCommand`) owns only flag syntax, duplicate/missing-flag detection, file reading, JSON parsing, root-shape validation, output presentation, and exit-code selection — every semantic rule (adequacy, compatibility, binding, coordinate mapping, tolerance, relationship) stays owned by the domain module, and `evaluateReferenceCandidateFidelityFromArtifactRoots` owns artifact-root reading/orchestration. `TOP_LEVEL_HELP` and `runCli`'s dispatch table were updated; `cli.ts`'s existing "no browser/persistence implementation of its own" regression test (`B5-TST-008,009`) continues to pass unchanged, confirming the new command imports neither `artifacts/` nor a filesystem-write function.
|
|
124
|
+
|
|
125
|
+
## CLI exit behavior
|
|
126
|
+
|
|
127
|
+
Mirrors `evaluate-contract`'s exact precedent: `--enforce` is applied only after the fidelity evaluation has already been computed, and affects only the process exit status for an already-final `state: 'fail'` result (never its printed content). `state: 'not-evaluated'` always exits `0` regardless of `--enforce` — a reference-adequacy or compatibility blocker is a successful, honest non-evaluation, never treated as a design mismatch or an execution error. An actual operational failure (invalid syntax, unreadable/malformed `--reference`/`--candidate`, malformed `--bindings-file`, or an invalid binding declaration) always exits nonzero regardless of `--enforce`. Verified end-to-end by three CLI tests reproducing pass/fail/not-evaluated with and without `--enforce`.
|
|
128
|
+
|
|
129
|
+
## Boundedness
|
|
130
|
+
|
|
131
|
+
Reuses existing bounds without introducing new ones: `reference.requirements` is already bounded to ≤50 (`MAX_REFERENCE_REQUIREMENTS`, Prompt 3) at construction time, so `requirementResults` is bounded identically (1:1 with requirements); `bindingDeclarations` is already bounded to ≤20 (`MAX_REFERENCE_RUNTIME_BINDINGS`, Prompt 5); `deriveLayoutRelationships`'s own `MAX_CONFIGURED_TARGETS_FOR_RELATIONSHIPS`/`MAX_PAIRWISE_RELATIONSHIP_RECORDS` bounds (v0.4) apply unchanged, since it is called exactly once per relationship requirement over the same candidate. No unbounded diff/diagnostic collection is ever produced — Prompt 6 evaluates only explicitly authored Prompt 3 requirements, never "all visible differences."
|
|
132
|
+
|
|
133
|
+
## Provenance
|
|
134
|
+
|
|
135
|
+
Each `ReferenceRequirementFidelityResult` carries `requirementId`, `category`, `expectedDependentMode` (when present), `subject` (which itself names the reference region(s)/property/relationship/measurement), `boundRuntimeTargets`, and either the numeric (`referenceValue`/`candidateRawValue`/`candidateValue`/`delta`/`tolerance`) or relationship (`expectedRelationship`/`actualRelationship`) evidence that produced its status. The top-level `ReferenceCandidateFidelityEvaluation` additionally carries `referenceId`/`referenceRequestId`/`candidateObservationId`/`candidateRequestId`, the full Prompt 3 `adequacy`, the full Prompt 4 `compatibility` (when computed), and the full Prompt 5 `bindings` evaluation (when computed) — nothing is an unexplained delta.
|
|
136
|
+
|
|
137
|
+
## Files changed
|
|
138
|
+
|
|
139
|
+
New:
|
|
140
|
+
- `src/domain/externalReferenceFidelity.ts`
|
|
141
|
+
- `src/application/referenceFidelityEvaluationService.ts`
|
|
142
|
+
- `tests/unit/externalReferenceFidelity.test.ts`
|
|
143
|
+
- `tests/unit/cliEvaluateReferenceFidelity.test.ts`
|
|
144
|
+
|
|
145
|
+
Modified:
|
|
146
|
+
- `src/cli.ts` (`evaluate-reference-fidelity` command: help text, argument parser, `loadBindingsFile`, run function, `TOP_LEVEL_HELP`, dispatch table)
|
|
147
|
+
- `src/index.ts` (public export surface for the above)
|
|
148
|
+
- `docs/ARCHITECTURE.md`, `docs/CONTRACTS.md`, `docs/WORKFLOWS.md`, `docs/COMMANDS.md`
|
|
149
|
+
|
|
150
|
+
## Tests
|
|
151
|
+
|
|
152
|
+
927 unit tests total (889 pre-existing + 38 new), all pure domain/application/CLI-level tests — zero Chromium launches for fidelity evaluation itself (the one CLI end-to-end test constructs its candidate `ObservationArtifact` by writing a `manifest.json` directly to disk, exactly as `tests/unit/artifactWriter.test.ts` already does for writer-focused tests, never via `observe()`/a real browser).
|
|
153
|
+
|
|
154
|
+
- `tests/unit/externalReferenceFidelity.test.ts` (29 tests): behaviors A–W from the task's own behavior model, including the exact worked example (424 reference px, 223→446 candidate px, delta 22, tolerance ±4, FAIL), the exact tolerance boundary (428 passes, 428.0001 fails), `exact`/`absolute-reference-px`/`percent` tolerance kinds, 2x coordinate scale, aspect-ratio-mismatch and missing-viewport coordinate-mapping unavailability, compatibility/adequacy blockers, bound/ambiguous/unavailable/undeclared bindings, hidden-target unavailability, relationship pass/fail/family-scoping and not-exhibited-by-reference, a supported numeric measurement (horizontal-gap), deterministic ordering and pure-function repeatability, all-pass/one-fail/one-unavailable overall results, zero-requirements inadequacy, full three-input immutability, fail-closed structural validation, absence of source-ownership fields, and a source-scan confirming no browser/filesystem import.
|
|
155
|
+
- `tests/unit/cliEvaluateReferenceFidelity.test.ts` (9 tests): help text, top-level help listing, required-flag errors, missing `--reference`/`--candidate` errors, malformed `--bindings-file` (bad JSON/wrong root shape/unknown field), and three full end-to-end runs (pass, fail, not-evaluated) each checked with and without `--enforce`, reproducing the task's exact worked example through the real CLI/application/domain stack.
|
|
156
|
+
|
|
157
|
+
## Validation results
|
|
158
|
+
|
|
159
|
+
All commands run from the repository root, after the implementation commit:
|
|
160
|
+
|
|
161
|
+
- `npm run typecheck` — pass, zero errors.
|
|
162
|
+
- `npm run lint` — pass, zero errors/warnings.
|
|
163
|
+
- `npm test` — 48 test files, 927 tests, all pass.
|
|
164
|
+
- `npm run build` — pass, clean `tsc` compile.
|
|
165
|
+
- `npm run check:docs` — pass (17 required files present, `ROADMAP.md` format intact — no implementation batches added).
|
|
166
|
+
- `git diff --check` — exit 0, no whitespace errors.
|
|
167
|
+
- `npm pack --dry-run` — pass; `dist/domain/externalReferenceFidelity.{js,d.ts,js.map}` and `dist/application/referenceFidelityEvaluationService.{js,d.ts,js.map}` confirmed present in the tarball listing.
|
|
168
|
+
- `npm run test:security` — pass (5 + 63 = 68 tests), run because a new CLI/config-file input surface (`--bindings-file`) was added.
|
|
169
|
+
- `npm run test:browser` — pass on re-run (9 files, 120 tests). One transient failure (`v0.2: an ambiguous locator stops immediately...`) occurred on the first full-suite run; re-running that single test in isolation passed immediately, and a full second `test:browser` run passed 120/120 — confirmed environmental flakiness in a real-Chromium timing-sensitive test unrelated to any file this prompt touched (Prompt 6 never imports or modifies `chromiumAdapter.ts`/`evidenceCapture.ts`), not a regression.
|
|
170
|
+
|
|
171
|
+
## Regression results
|
|
172
|
+
|
|
173
|
+
- Prompt 1 external-reference foundation: unaffected — `externalReference.ts` untouched (only its already-exported type guards were imported).
|
|
174
|
+
- Prompt 2 regions/shared relationships: unaffected — `externalReferenceRegions.ts`/`externalReferenceRegionRelationships.ts` untouched.
|
|
175
|
+
- Prompt 3 requirements/tolerances/adequacy: unaffected — `externalReferenceRequirements.ts` untouched; its exported functions were called, never modified.
|
|
176
|
+
- Prompt 4 compatibility: unaffected — `externalReferenceCompatibility.ts`/`explicitState.ts`/`externalReferenceApplicability.ts` untouched.
|
|
177
|
+
- Prompt 5 binding: unaffected — `externalReferenceRuntimeBinding.ts` untouched.
|
|
178
|
+
- v0.4 comparison/relationship behavior: unaffected — `comparisonEngine.ts`/`relationships.ts` untouched (only already-exported functions were called).
|
|
179
|
+
- v0.5 contracts/evaluation: unaffected — no file in `frontendContracts*.ts` touched.
|
|
180
|
+
- v0.6 bounded context/correlation: unaffected — `boundedAgentContext*.ts` untouched; read for precedent only in an earlier prompt, not this one.
|
|
181
|
+
- Full unit (927/927), browser (120/120 on the confirming re-run), and security (68/68) suites all pass with zero regressions.
|
|
182
|
+
|
|
183
|
+
## Security impact
|
|
184
|
+
|
|
185
|
+
- No new external input surface beyond a local, caller-controlled JSON file (`--bindings-file`) — the same trust boundary as every existing `--*-file` flag.
|
|
186
|
+
- `evaluateReferenceCandidateFidelity` performs no filesystem access, no network access, and no browser/Chromium invocation of any kind; `evaluateReferenceCandidateFidelityFromArtifactRoots` performs only read-only artifact-manifest reads through the existing readers.
|
|
187
|
+
- `test:security` (policy + real-Chromium adapter tests) re-run and passing, confirming no regression to the existing safety/navigation policy surface (untouched by this prompt).
|
|
188
|
+
|
|
189
|
+
## Documentation changes
|
|
190
|
+
|
|
191
|
+
- `docs/CONTRACTS.md` — new "v0.7 Prompt 6 structured reference-vs-candidate fidelity evaluation" section (full type shapes, evaluation order, coordinate-mapping/aspect-ratio rules, tolerance reuse, region-property/measurement/relationship evaluation, binding gate, category preservation, overall-result rule, persistence decision, CLI summary).
|
|
192
|
+
- `docs/ARCHITECTURE.md` — new paragraph in the "Planned v0.7–v0.10" section describing the Prompt 6 module additions and their reuse of Prompts 3/4/5 and v0.4.
|
|
193
|
+
- `docs/WORKFLOWS.md` — "Current external-reference foundation workflow" retitled to "Prompts 1-6" and extended with the `evaluate-reference-fidelity` command flow and test-coverage summary.
|
|
194
|
+
- `docs/COMMANDS.md` — new `## `evaluate-reference-fidelity`` section, mirroring the existing `## `evaluate-contract`` section's structure exactly.
|
|
195
|
+
|
|
196
|
+
## Tooling incidents
|
|
197
|
+
|
|
198
|
+
None. No orchestrator was invoked (direct-implementation mode used throughout, consistent with Prompts 2–6); no background/speculative subagent writes occurred; the Prompt 1 stray-fork-writes stash remains untouched, unapplied, and unmined as precedent. The one transient real-Chromium test failure encountered during validation is documented under Validation/Regression results above as confirmed flakiness, not a tooling or implementation defect.
|
|
199
|
+
|
|
200
|
+
## Out-of-scope confirmation
|
|
201
|
+
|
|
202
|
+
This prompt implements no bounded coding-agent correction packet, no v0.6 bounded-agent-context change, no source/static correlation change, no coding-agent invocation, no source editing, no rerender/retry loop, no baseline/per-change overall workflow composition, no viewer, no annotation, no automatic target matching, no pixel/image similarity (no SSIM/perceptual hash/OCR/computer vision of any kind), and no new style-fidelity family (background color, text color, font size/weight, line height, border radius, shadow, gradient, opacity, icons) beyond Prompt 3's already-frozen requirement vocabulary. `evaluateReferenceCandidateFidelity` never attaches `sourceOwner`/`sourceFile`/`component`/`symbol`/`causedBy` to any result (verified by a dedicated test scanning the serialized evaluation for those exact terms).
|
|
203
|
+
|
|
204
|
+
## Known limitations
|
|
205
|
+
|
|
206
|
+
- The coordinate-mapping model supports only a single, deliberately bounded full-frame scale per evaluation — no per-region cropping, offset, rotation, or perspective mapping exists or was attempted, per the task's own explicit prohibition.
|
|
207
|
+
- `expectedDependentMode` (`required`/`permitted`) is carried through as provenance only and never changes the pass/fail rule, since Prompt 3 itself never implemented a directional evaluation difference for its own expectation/adequacy derivation (unlike v0.5's runtime-directional contract clauses) — documented explicitly rather than silently ignored.
|
|
208
|
+
- No `--output`/persistence flag exists for `evaluate-reference-fidelity`; a caller needing to save a fidelity result must do so itself (e.g. redirecting stdout, or consuming the programmatic API directly) until/unless a later prompt's architecture proves in-repository persistence necessary.
|
|
209
|
+
|
|
210
|
+
## Remaining risks
|
|
211
|
+
|
|
212
|
+
- None identified that block this prompt's own scope. The primary forward consideration for Prompt 7 (bounded visual-fidelity mismatch projection and v0.6 bounded-agent-context integration) is how it will attach `ReferenceRequirementFidelityResult`/`ReferenceCandidateFidelityEvaluation` evidence to `RuntimeStaticCorrelationRecord`-shaped context without conflating the two identity domains this prompt was careful to keep separate (reference-region/runtime-target vs. runtime-target/static-source) — flagged for that prompt's own precedent review, not preempted here.
|
|
213
|
+
|
|
214
|
+
## Exact next action
|
|
215
|
+
|
|
216
|
+
v0.7 Prompt 7 — bounded visual-fidelity mismatch projection and v0.6 bounded-agent-context integration.
|