@keboola/validate-ui 0.4.0 → 0.5.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md ADDED
@@ -0,0 +1,13 @@
1
+ # @keboola/validate-ui — agent guide
2
+
3
+ Closed-loop validator for generated UI. Static gates check code; this checks the
4
+ rendered result. Reach for it after generating/scaffolding a module, before a PR.
5
+
6
+ - **Run it:** `validate-ui --url <url> --brief "<the brief>"` (or `--serve <distDir> --route <path>`). Exit 0 = all axes pass. `--json` for machine output.
7
+ - **Programmatic:** `validate({ url, brief })` → `AggregateVerdict`. Building a pipeline? `capture()` for artifacts, `runAxes(axes, ctx)` to run a subset.
8
+ - **Five axes** live one-per-file in `src/axes/` (`runtime-health`, `a11y`, `visual-brand`, `brief-conformance`, `compare`), each implementing the `Axis` contract in `src/types.ts`. Add an axis: implement `{ name, run }`, register it in `src/axes/index.ts`.
9
+ - **Contract boundary:** everything an axis needs is on `AxisContext` (`artifact`, live `page`, `brief`, `baseline`, `captureUnderBrand`, `comparison`, `compareConfig`). Don't reach outside it.
10
+ - **Compare (old-vs-new)** is semantic, not a value multiset — `src/compare/`. Both sides reduce to a serializable `CompareSnapshot` (regions with key/state/values) via `extractRegions`; the pure `diffSnapshots` localizes every delta to a region, suppresses declared `expectedAbsent` deltas (recorded, not dropped), and reports an empty/error `region-degraded` instead of value loss. Multi-route: `validate-ui compare --new <url> (--old <url> | --old-snapshots <dir>) --routes a,b,c [--config f.json]` — `dedupeSharedChrome` collapses a delta seen on ≥2 routes into one finding so shared chrome isn't multiplied. The `compare` axis only activates when `context.comparison` is set, so single-page runs are unaffected.
11
+ - **Do not** depend on `@keboola/e2e-testing` (private) — the screenshot-prep helper is reimplemented in `src/prepare-screenshot.ts`.
12
+ - **Want one pass/fail number + a gate exit code**, not raw findings? Use `@keboola/ui-gen-bench` — `ui-gen-bench eval --serve <dir> --route <path> --brief "…"` boots via this validator, scores the verdict, and exits 0/1. That's the per-generation gate; this package is the validator underneath it.
13
+ - **VLM backend** for `brief-conformance` is chosen in `src/axes/vlm-provider.ts`: Keboola LLM proxy (base URL `VALIDATE_UI_LLM_BASE_URL` → `ANTHROPIC_BASE_URL`; token `VALIDATE_UI_LLM_TOKEN` → `KBC_TOKEN` → `ANTHROPIC_API_KEY` — so kai-agent's SDK-native `ANTHROPIC_BASE_URL` + `ANTHROPIC_API_KEY` pair works, preferred) → raw `ANTHROPIC_API_KEY` (no proxy base URL) → skip. Route new VLM calls through `resolveVlmProvider()`, never `new Anthropic()` directly.
package/CHANGELOG.md ADDED
@@ -0,0 +1,116 @@
1
+ # @keboola/validate-ui
2
+
3
+ ## 0.5.2
4
+
5
+ ### Patch Changes
6
+
7
+ - The published tarball now carries a `CHANGELOG.md`, so a consumer can read what changed in a release — breaking changes included — without leaving their `node_modules`.
8
+
9
+ It could not simply be added to `files`. This repo is private, so every reference `@changesets/changelog-github` emits is a dead link for anyone reading from npm: PR links, commit SHAs, author handles, and the Linear and cross-repo links that changeset prose carries. Across the publishable packages that came to 490 PR links, 863 commit links, 490 author credits and 102 dependency-bump blocks.
10
+
11
+ The published file is generated, not maintained. `scripts/public-changelog.mjs` removes those links ahead of the publish and leaves the prose. An identifier the author typed themselves stays as text — `UT-4009`, `connection#8040` — because it is part of the sentence and, without its URL, resolves to nothing outside Keboola. The repo-side `CHANGELOG.md` keeps every link, because that is how a release gets traced internally. It is rewritten only for the moment the tarballs are packed, then restored — which the release also depends on: `changesets/action` reads each changelog back off disk _after_ the publish command returns, to build that version's GitHub Release body, so without the restore the internal releases would carry the public text. A reference the rules do not cover fails the publish rather than shipping.
12
+
13
+ One incidental fix: a hex colour written as `#222529` in changeset prose had been autolinked into a link to issue 222529. Unwrapping restores the colour, so the published notes read as the author wrote them.
14
+
15
+ - Updated dependencies:
16
+ - @keboola/brand-registry@1.9.3
17
+
18
+ ## 0.5.1
19
+
20
+ ### Patch Changes
21
+
22
+ - Build with tsdown (rolldown) instead of the now-unmaintained tsup. Output layout, exports map, and shipped declarations are unchanged (attw-verified per package); chunk byte sizes shift slightly with rolldown's codegen, and `@keboola/design`'s size budgets are trued up to the new measurements (largest delta: the main bundle cap moves 320→330 KB). `@keboola/design` builds with `platform: 'browser'`, so bundled CJS deps (react-dropzone, react-day-picker) resolve to their ESM builds instead of a UMD `main` that would drag a Node-only runtime require into browser loads.
23
+
24
+ - Updated dependencies:
25
+ - @keboola/brand-registry@1.9.2
26
+
27
+ ## 0.5.0
28
+
29
+ ### Minor Changes
30
+
31
+ - The `a11y` axis now reports content that is silently cut off — text or an input value wider than the box holding it, with no way for the user to reveal the rest.
32
+
33
+ This closes a real hole. A page can be free of console errors, axe violations and layout overflow while still showing the user the wrong thing: an editable financial grid whose cells were narrower than their own padding rendered a value of `800` as `8`, and scored **100** on every axis. Clipping inside an element is not a runtime error, not an axe rule, and not page-level overflow, so nothing saw it.
34
+
35
+ Deliberate truncation is not reported. `text-overflow: ellipsis` announces itself, a scrollable box (`overflow: auto | scroll`) still reaches its content, and visually-hidden text (`sr-only`) is a 1px box holding a whole sentence by design — flagging that would punish the accessible thing to do. Only silent loss is a finding: `overflow: hidden | clip` without an ellipsis, and inputs, which cannot be scrolled by eye.
36
+
37
+ Findings name the element by accessible name or id where possible rather than by utility classes (which identify every cell in a grid identically), quantify the gap, and are capped at 10 so one broken component cannot flood the report. When an ancestor does the clipping it is reported once, rather than once per descendant.
38
+
39
+ ### Patch Changes
40
+
41
+ - Updated dependencies:
42
+ - @keboola/brand-registry@1.8.0
43
+
44
+ ## 0.4.1
45
+
46
+ ### Patch Changes
47
+
48
+ - These packages now ship their `AGENTS.md` usage contract to npm, so external
49
+ consumers (and their AI agents) can read it at
50
+ `node_modules/@keboola/<name>/AGENTS.md`, version-pinned to the release they
51
+ actually installed.
52
+
53
+ Previously only `@keboola/design` published its `AGENTS.md`; every other package
54
+ omitted it from `files`, so instructions that point agents at that path — such as
55
+ `apps/boilerplate/AGENTS.md` — silently resolved to nothing outside the monorepo.
56
+ No code or type changes.
57
+
58
+ - Fix the `visual-brand` axis reporting "chrome did not re-skin" for apps that re-skin correctly.
59
+
60
+ The alt-brand comparison reused the baseline-drift pixel threshold (`pixelmatch` `threshold: 0.1`). That is deliberately coarse so run-to-run antialiasing does not register as drift — but a brand re-skin is mostly broad, low-amplitude hue and tone shifts across backgrounds, muted surfaces and borders, which the coarse threshold discards almost entirely.
61
+
62
+ Measured on the reference app, an obvious and correct re-skin — new logo, purple accents, different typeface, 72% of pixels changed by raw comparison — scored 0.0008 on `/` and 0.0038 on `/tokens`, both **under** the 0.005 re-skin floor. So a correctly branded app was told its chrome did not re-skin: a false negative that penalises correct code, the worst failure mode for a gate.
63
+
64
+ The two comparisons now carry their own sensitivities: drift detection stays at `0.1`, the re-skin check uses `0.02`. The same pages now score 0.008 and 0.012. Baseline-drift behaviour is unchanged.
65
+
66
+ - The `visual-brand` alt-brand check no longer fails apps that cannot switch brands at runtime.
67
+
68
+ `captureUnderBrand` requests the alternate brand with `?brand=<id>`, which only works if the app reads that param and re-resolves its `ThemeProvider`. Most apps deliberately don't: the brand is build configuration, and a query-param brand switcher is not something a scaffold should impose on every generated app. Those apps captured the same brand twice, and the identical screenshots were reported as `chrome did not re-skin` — a **serious** finding against code that is entirely correct.
69
+
70
+ The check now confirms the requested brand actually became active (via `data-brand` on the root element) before comparing. If it did not, the finding is a **minor** "not assessed" note explaining how to opt in, rather than a failure.
71
+
72
+ Note the consequence: for an app that does not honour `?brand=`, `visual-brand`'s only finding is that skip, so the axis is reported as not-applicable and excluded from the total — its overflow and interactive-contrast checks stop contributing too. Apps that want the axis scored should register the alternate brand and resolve the active id from the `brand` query parameter.
73
+
74
+ - Updated dependencies:
75
+ - @keboola/brand-registry@1.7.1
76
+
77
+ ## 0.4.0
78
+
79
+ ### Minor Changes
80
+
81
+ - Guarantee AA-safe interactive brand tokens instead of leaving contrast to the consumer.
82
+ - `@keboola/brand-registry` exports `checkInteractiveContrast(brand)` (plus `contrastRatio`, `AA_CONTRAST_MIN`, the `ContrastWarning` type, and a `brands` array): a non-fatal check that reports interactive foreground↔background pairings — button/link labels on branded fills, and `primary`-on-`background` — that fall below WCAG-AA (4.5:1), across both the light and dark palettes.
83
+ - `@keboola/validate-ui`'s visual-brand axis now asserts the alternate brand's interactive-token contrast (non-fatal `moderate` findings), not just that chrome re-skins.
84
+ - The `keboola` V2 brand's light `primary` is AA-tuned to blue-600 (`37 99 235`, 5.17:1 white-on-fill), kept in lock-step with `@keboola/design/styles.css`; the dark palette and both `keboola-legacy` brands are unchanged.
85
+
86
+ ### Patch Changes
87
+
88
+ - Updated dependencies:
89
+ - @keboola/brand-registry@1.5.0
90
+
91
+ ## 0.3.0
92
+
93
+ ### Minor Changes
94
+
95
+ - Add a semantic old-vs-new `compare` axis. Instead of a naive numeric multiset diff, both sides reduce to a serializable `CompareSnapshot` (regions keyed by `data-region`/testid/landmark/label/heading, each with a backend `state` and normalized value tokens). The pure `diffSnapshots` localizes every delta to a region: a removed region is `serious` unless matched by an `expectedAbsent` allowlist rule (suppressed with its reason, never dropped), a region that degraded to an empty/error state is reported informationally rather than as value loss, and surviving regions get a per-region value diff. A multi-route runner (`compareRoutes`, `validate-ui compare --new … --old … --routes …`) collapses a delta seen on ≥2 routes into one shared-chrome finding so shared elements aren't multiplied across pages. The axis only activates when a comparison snapshot is supplied, so existing single-page runs are unaffected.
96
+
97
+ ## 0.2.0
98
+
99
+ ### Minor Changes
100
+
101
+ - feat: publish @keboola/validate-ui and @keboola/ui-gen-bench to npm (public)
102
+
103
+ First public release. Both drop `private` and gain `publishConfig.access: public`
104
+ so external consumers can install them and run the per-generation `eval`
105
+ self-check against generated Keboola UI. `@keboola/validate-ui-mcp` stays private
106
+ for now (added to the changeset `ignore` list). The `workspace:^` interdep
107
+ (`ui-gen-bench` → `validate-ui`) is rewritten to the published version by
108
+ `pnpm publish -r` at release time.
109
+
110
+ - brief-conformance can now call the Claude vision model through Keboola's LLM
111
+ proxy, not only a raw `ANTHROPIC_API_KEY`. The VLM call is behind a provider
112
+ abstraction (`src/axes/vlm-provider.ts`) with two env-selected backends: the
113
+ Keboola LLM path (`VALIDATE_UI_LLM_BASE_URL` + `VALIDATE_UI_LLM_TOKEN`,
114
+ mirroring how kai-agent routes Claude, preferred for team plans with no
115
+ dedicated key) and the existing raw-key path. The axis still skips gracefully
116
+ when neither is configured.