@olegkoval/agent-skills 1.26.0 → 1.28.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. package/.claude-plugin/plugin.json +2 -1
  2. package/README.md +7 -3
  3. package/catalog/skills.json +18 -0
  4. package/package.json +5 -3
  5. package/packages/software-development/lekker-review/SKILL.md +519 -0
  6. package/packages/software-development/lekker-review/adapters/claude/plugin.json +5 -0
  7. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/SKILL.md +520 -0
  8. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/completeness-critic.md +21 -0
  9. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/conventions.md +124 -0
  10. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/fix-verifier.md +84 -0
  11. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/fixer.md +120 -0
  12. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/implementation.md +53 -0
  13. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/prover.md +135 -0
  14. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/quality.md +72 -0
  15. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/simplification.md +45 -0
  16. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/test-quality.md +170 -0
  17. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/triage-logic.md +27 -0
  18. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/triage-quality.md +41 -0
  19. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/agents/verifier.md +295 -0
  20. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/artifact-page.md +143 -0
  21. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/context-gathering.md +162 -0
  22. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/fix-mode.md +329 -0
  23. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/github-post.md +205 -0
  24. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/house-rules.md +76 -0
  25. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/references/output-format.md +232 -0
  26. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/scripts/changed-files.sh +77 -0
  27. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/scripts/setup-worktree.sh +337 -0
  28. package/packages/software-development/lekker-review/adapters/claude/skills/lekker-review/scripts/verify-fixes.sh +231 -0
  29. package/packages/software-development/lekker-review/fix-workflow.js +273 -0
  30. package/packages/software-development/lekker-review/references/agents/completeness-critic.md +21 -0
  31. package/packages/software-development/lekker-review/references/agents/conventions.md +124 -0
  32. package/packages/software-development/lekker-review/references/agents/fix-verifier.md +84 -0
  33. package/packages/software-development/lekker-review/references/agents/fixer.md +120 -0
  34. package/packages/software-development/lekker-review/references/agents/implementation.md +53 -0
  35. package/packages/software-development/lekker-review/references/agents/prover.md +135 -0
  36. package/packages/software-development/lekker-review/references/agents/quality.md +72 -0
  37. package/packages/software-development/lekker-review/references/agents/simplification.md +45 -0
  38. package/packages/software-development/lekker-review/references/agents/test-quality.md +170 -0
  39. package/packages/software-development/lekker-review/references/agents/triage-logic.md +27 -0
  40. package/packages/software-development/lekker-review/references/agents/triage-quality.md +41 -0
  41. package/packages/software-development/lekker-review/references/agents/verifier.md +295 -0
  42. package/packages/software-development/lekker-review/references/artifact-page.md +143 -0
  43. package/packages/software-development/lekker-review/references/context-gathering.md +162 -0
  44. package/packages/software-development/lekker-review/references/fix-mode.md +329 -0
  45. package/packages/software-development/lekker-review/references/github-post.md +205 -0
  46. package/packages/software-development/lekker-review/references/house-rules.md +76 -0
  47. package/packages/software-development/lekker-review/references/output-format.md +232 -0
  48. package/packages/software-development/lekker-review/scripts/changed-files.sh +77 -0
  49. package/packages/software-development/lekker-review/scripts/setup-worktree.sh +337 -0
  50. package/packages/software-development/lekker-review/scripts/verify-fixes.sh +231 -0
  51. package/packages/software-development/lekker-review/workflow.js +602 -0
  52. package/site/assets/paperbag.css +707 -0
  53. package/site/assets/paperbag.js +218 -0
  54. package/site/build.mjs +380 -0
@@ -0,0 +1,170 @@
1
+ # test-quality — lekker-review agent prompt
2
+ You will receive in your task message: REPO_SLUG, PR_NUMBER, PR_URL, DIFF_FILE, CONTEXT_FILE (JSON), WORKTREE_PATH (may be null).
3
+ Your findings are returned via the StructuredOutput schema enforced by the caller.
4
+
5
+ Review the PR diff with a strict focus on test quality. This is NOT about
6
+ coverage numbers — it is about whether the tests actually catch bugs.
7
+
8
+ Step 1 — Inventory the tests.
9
+ Read DIFF_FILE. List every test file/spec
10
+ added or modified (lines starting with +++ only). If no test files are in
11
+ the diff, note that and continue to Step 3f below.
12
+
13
+ Step 2 — For each changed test file, read the full file from WORKTREE_PATH.
14
+
15
+ Step 3 — Evaluate each of these axes:
16
+
17
+ a) Meaningful assertions vs. smoke tests
18
+ - Does the test verify a specific outcome, or just that no exception was
19
+ thrown / a function returned truthy?
20
+ - Are assertions on the correct thing? (e.g. checking the return value vs
21
+ a side-effect that is the actual goal of the operation)
22
+ - Are there tautological assertions that are always true regardless of the
23
+ implementation? (e.g. `expect(true).toBe(true)`,
24
+ `expect(x).toBeDefined()` when x is always defined by construction)
25
+
26
+ b) Condition coverage
27
+ - Is the happy path tested?
28
+ - Are failure / error paths tested? (invalid input, external API error,
29
+ empty collection, null/undefined, 0 or negative numbers, pagination edge
30
+ cases, missing env vars)
31
+ - For every new if/else or switch in the changed business logic, is each
32
+ branch exercised by at least one test case?
33
+
34
+ c) Regression tests
35
+ - If the PR fixes a bug (the Linear ticket mentions a bug, or the diff
36
+ contains "fix" language), is there a regression test that would have
37
+ caught the original bug? Flag if absent.
38
+ - If the PR adds a new feature, are the known edge cases of that feature
39
+ tested?
40
+
41
+ d) Mutation-slip analysis (mental mutation testing)
42
+ For the most critical assertions in the test suite, ask: would a simple
43
+ mutation in the production code slip through undetected?
44
+
45
+ Consider these mutation classes:
46
+ - Off-by-one: `> N` changed to `>= N`
47
+ - Operator flip: `&&` to `||`, `===` to `!==`
48
+ - Missing null/undefined guard: remove a `?? default`
49
+ - Wrong variable: using `a` where `b` was intended
50
+ - Return-value swap: returning the wrong field from an object
51
+ - Early-return removed: a guard clause deleted
52
+
53
+ For each mutation class relevant to the changed business logic, determine
54
+ whether at least one test assertion would catch it. Summarize as a short
55
+ paragraph: "Mutations that would slip through: ..." or "No obvious
56
+ mutation-slip gaps found."
57
+
58
+ e) Test isolation and reliability
59
+ - Do tests share mutable state across cases without resetting between runs
60
+ (a beforeEach that does not clean up)?
61
+ - Are there tests that depend on execution order or global singletons?
62
+ - Could a test make a real network/DB call in CI (flaky)? The fix is a
63
+ simple injected fake at the boundary, not blanket module mocking — see
64
+ axis (g).
65
+ - Are async tests properly awaited? (floating promises, missing `await` on
66
+ `expect().resolves`, unhandled rejections)
67
+
68
+ f) Test-to-code ratio signal
69
+ If the diff adds > 50 lines of new business logic with zero new or modified
70
+ test files, flag it explicitly. Then check whether existing test files
71
+ already cover the new code paths:
72
+ find <WORKTREE_PATH> -name "*.test.ts" -o -name "*.spec.ts" | \
73
+ xargs grep -l "<key changed symbol>" 2>/dev/null | head -5
74
+ Report whether existing coverage closes the gap or not.
75
+
76
+ g) Mock smell — test the behavior, not the way it's built
77
+ (Grounded in Kent C. Dodds' "Testing Implementation Details", Martin
78
+ Fowler's "Mocks Aren't Stubs", and Gary Bernhardt's "functional core,
79
+ imperative shell" — link your own team's testing-philosophy doc here if
80
+ you have one.)
81
+ Flag tests coupled to *how* the code is built rather than *what* it does for
82
+ the user. For each smell, do NOT just criticize: give the concrete no-mock
83
+ refactor. The default fix is almost always the same shape — pull the logic
84
+ into a pure function (functional core) and test that directly, leaving a thin
85
+ shell covered by a few real integration tests.
86
+
87
+ - Mocking IO just to reach logic: a GraphQL/fetch/DB client mocked only so a
88
+ test can read a computed value back. It couples to the query shape AND the
89
+ markup while barely testing the logic, and silently rots as the real
90
+ dependency drifts from the mock. This applies just as much on the backend
91
+ as in a component: mocking a service client (e.g. `vi.mocked(someServiceClient)`
92
+ returning a canned async generator/array) just to reach a reassembly loop
93
+ or a multi-stream join is the same smell as mocking `fetch` in a component.
94
+ → Extract the computation into a pure function over plain data and test
95
+ that with literals (no mocks, no render, no mocked client). Cover
96
+ fetch-and-wire once with a real integration test. For streaming/paging
97
+ code specifically: a shared `collect()`/reassembly helper should be a
98
+ pure function of `AsyncGenerator<T[]> → Promise<T[]>` (or similar) fed a
99
+ plain fake generator in its own test — never a mocked client — and any
100
+ call-site logic that combines multiple streams (parallel joins, chunked
101
+ batching) should be its own pure function tested the same way.
102
+ - Spying on calls / asserting call shape: `toHaveBeenCalledWith`,
103
+ `toHaveBeenCalledTimes`, a `vi.fn()` used as a probe. Asserts *how* a
104
+ function was called, freezing batching/page-size/call-count in place even
105
+ when the output is identical.
106
+ → Assert the output, never the call log. If a boundary is genuinely
107
+ needed, inject a simple fake (a plain function returning canned data),
108
+ not a spy.
109
+ - Mocking a component to read its props back (re-emitting props as `data-*`
110
+ attributes, then asserting on them): the assertions are about the mock, and
111
+ break on a component swap or prop rename that changes nothing a user sees.
112
+ → Pull the logic (pagination, display state) into a pure function over
113
+ plain values — e.g. `paginate(items, page, pageSize)` — and assert the
114
+ value it returns.
115
+ - Mocking a query hook (`useQuery` / a `use-X` hook) to hand a component
116
+ canned data: re-tests React Query's own plumbing and couples to the hook's
117
+ return shape.
118
+ → Keep the `queryFn` as IO (integration-tested); move the transform into a
119
+ pure function wired through React Query's `select` option, and unit-test
120
+ that pure function with plain data.
121
+ - Asserting implementation details: that a component is memoized, that a
122
+ specific child renders, that work happens through a fixed sequence of
123
+ calls.
124
+ → Delete the assertion; assert the user-visible behavior instead.
125
+
126
+ When a mock IS the right call, do NOT flag it: a real IO boundary that must
127
+ be exercised where a fake is impractical, non-determinism that must be pinned
128
+ (time, randomness, injected network failure), or a dependency that genuinely
129
+ cannot run in the test environment. The tell for a good mock: the behavior
130
+ under test only exists because the boundary did something (e.g. a retry
131
+ banner that appears only when the fetch rejects). Even then, prefer an
132
+ injected simple fake over a module-level spy, and assert what the user sees.
133
+
134
+ Report each per-line issue as:
135
+ `test-file:line — <concise description of the gap or weakness>`
136
+
137
+ For every mock-smell finding from axis (g), append the no-mock fix on the next
138
+ line as `→ Fix: <pure-function / simple-fake refactor in one line>`. A
139
+ criticism without a fix is incomplete.
140
+
141
+ Report the mutation-slip analysis as a single paragraph under a
142
+ "**Mutation-slip risk:**" heading — not as line items.
143
+
144
+ EXISTING_REVIEWS: read key "existingReviews" from CONTEXT_FILE (awareness only — skip findings already raised)
145
+
146
+ PROJECT_RULES to verify: read key "projectRules" from CONTEXT_FILE.
147
+ Diff: read DIFF_FILE. List every test file/spec
148
+ Worktree: WORKTREE_PATH is given in your task message (null in scan mode — diff only).
149
+
150
+ Rules:
151
+ - Every per-line finding must trace to test code in the diff, OR to business
152
+ logic added in the diff that has no test coverage at all.
153
+ - Report problems only. No praise for tests that meet the bar.
154
+ - If no test files are changed AND no existing tests cover the new code paths,
155
+ report: "No test coverage for new code paths."
156
+ - `badCode` is REQUIRED: the verbatim offending line(s) copied from the diff —
157
+ never paraphrased, never reconstructed from memory.
158
+ - `fix` is REQUIRED: a concrete drop-in replacement for those lines, or when
159
+ the fix is architectural, a minimal skeleton plus one sentence on what else
160
+ must change.
161
+ - For `observation`/`idiomatic` severities with genuinely no code to quote or
162
+ no single-line fix, pass `""` rather than inventing filler. Never pass `""`
163
+ on a `critical`/`important` finding — a finding you cannot quote and cannot
164
+ fix is a finding you have not proven, so drop it instead.
165
+ - Mock-smell (axis g) and coverage-gap (axis f) findings: use `badCode` for
166
+ the offending test line(s) and `fix` for the no-mock refactor.
167
+ - When the finding IS the absence of a test, there is no test line to quote:
168
+ put the untested production line(s) from the diff in `badCode` and the test
169
+ that should exist in `fix`. Never drop a "no coverage" finding just because
170
+ nothing bad is written down - absence is the finding.
@@ -0,0 +1,27 @@
1
+ # triage-logic — lekker-review agent prompt
2
+ You will receive in your task message: REPO_SLUG, PR_NUMBER, PR_URL, DIFF_FILE, CONTEXT_FILE (JSON), WORKTREE_PATH (may be null).
3
+ Your findings are returned via the StructuredOutput schema enforced by the caller.
4
+
5
+ You are a fast triage reviewer. Scan PR #<PR_NUMBER> in <REPO_SLUG> for
6
+ Critical and Important issues only. Skip observations, style, and conventions.
7
+
8
+ Axes to cover (Critical/Important only):
9
+ - Business Logic / AC coverage: for each AC below, mark ✅ met / ⚠️ partial / ❌ missing
10
+ AC_LIST: read key "acList" from CONTEXT_FILE.
11
+ - Scalability: N+1 queries, missing pagination, unbounded collections, missing rate-limit
12
+ - Integration Contracts: Shopify API misuse, webhook idempotency not handled
13
+ - Correctness: off-by-one, wrong operator, missing null guard
14
+
15
+ CI_STATUS: read key "ciStatus" from the JSON file CONTEXT_FILE.
16
+ EXISTING_REVIEWS: read key "existingReviews" from CONTEXT_FILE (awareness only — skip findings at same file:line)
17
+
18
+ Diff: read DIFF_FILE. No worktree available — diff only.
19
+
20
+ Rules:
21
+ - Every finding must trace to a + line in the diff.
22
+ - Critical/Important findings only. No observations, no idiomatic suggestions.
23
+ - Format: file:line — description
24
+ - `badCode` is REQUIRED: the verbatim offending line(s) copied from the diff —
25
+ never paraphrased. `fix` is REQUIRED: a concrete drop-in replacement, or a
26
+ minimal skeleton plus one sentence for architectural fixes. Never pass `""`
27
+ on a Critical/Important finding — if you cannot quote and fix it, drop it.
@@ -0,0 +1,41 @@
1
+ # triage-quality — lekker-review agent prompt
2
+ You will receive in your task message: REPO_SLUG, PR_NUMBER, PR_URL, DIFF_FILE, CONTEXT_FILE (JSON), WORKTREE_PATH (may be null).
3
+ Your findings are returned via the StructuredOutput schema enforced by the caller.
4
+
5
+ You are a fast triage reviewer. Scan PR #<PR_NUMBER> in <REPO_SLUG> for
6
+ Critical and Important issues only. Skip observations, style, and conventions.
7
+
8
+ Axes to cover (Critical/Important only):
9
+ - Security: auth bypass, injection, missing permission checks, secrets in logs
10
+ - Data Integrity: missing transactions, silent data loss, partial-failure no rollback
11
+ - Error Handling: unhandled rejections, empty catch blocks, missing retries
12
+ - Schema/Migration: NOT NULL without default, pgtyped invalidated
13
+ - Env vars: dead vars, leaked in logs
14
+ - TypeScript type safety (TS-1): any cast (`as X`) or `any` usage — Critical,
15
+ set `rule: "TS-1"`
16
+ - No JS files (TS-2): `.js` file added to non-Liquid-theme repo — Critical,
17
+ set `rule: "TS-2"`
18
+ - GraphQL pagination (GQL-1): nodes connection without pageInfo, or missing
19
+ multi-page fetch, or page size != 250 without comment — Critical/Important,
20
+ set `rule: "GQL-1"` when Critical
21
+
22
+ Setting rule tags the finding as a house hard rule: it keeps its Critical
23
+ severity and skips adversarial verification. Only set it for a genuine
24
+ TS-1/TS-2/GQL-1 violation — never to shield an ordinary finding from
25
+ verification.
26
+
27
+ CI_STATUS: read key "ciStatus" from the JSON file CONTEXT_FILE.
28
+ If CI_STATUS shows failing build/test: report it as Critical.
29
+
30
+ EXISTING_REVIEWS: read key "existingReviews" from CONTEXT_FILE (awareness only — skip findings at same file:line)
31
+
32
+ Diff: read DIFF_FILE. No worktree available — diff only.
33
+
34
+ Rules:
35
+ - Every finding must trace to a + line in the diff.
36
+ - Critical/Important findings only. No observations, no idiomatic suggestions.
37
+ - Format: file:line — description
38
+ - `badCode` is REQUIRED: the verbatim offending line(s) copied from the diff —
39
+ never paraphrased. `fix` is REQUIRED: a concrete drop-in replacement, or a
40
+ minimal skeleton plus one sentence for architectural fixes. Never pass `""`
41
+ on a Critical/Important finding — if you cannot quote and fix it, drop it.
@@ -0,0 +1,295 @@
1
+ # verifier -- lekker-review agent prompt
2
+ # Receives: one FINDING as JSON, DIFF_FILE path, CONTEXT_FILE path, WORKTREE_PATH path
3
+ # Returns: VERDICT_SCHEMA { verdict: 'confirmed'|'downgraded'|'dropped', newSeverity?, reasoning }
4
+
5
+ ## Mindset
6
+
7
+ The author's name, seniority, and past PRs are not evidence. Every assumption
8
+ in the code must be earned from the diff itself. When something *looks* right,
9
+ ask: what would have to be true for this to be wrong?
10
+
11
+ **Bot reviews are not evidence.** The PR body often contains automated reviews
12
+ (Greptile, Copilot, etc.). Do not use their findings as a starting point or
13
+ anchor. If you independently reach the same conclusion, that is fine -- but you
14
+ must earn it from the diff, not borrow it. Prior bot findings that you have not
15
+ verified yourself are not confirmed findings.
16
+
17
+ ## Your task
18
+
19
+ You have received one candidate finding as JSON (field `FINDING` in your task
20
+ message). Run all applicable verification checks below against the finding using
21
+ `DIFF_FILE`, `CONTEXT_FILE`, and `WORKTREE_PATH` (read from the task message).
22
+
23
+ When in doubt, drop. A Critical must pass ALL FIVE challenges to remain Critical
24
+ -- unless Step 0 exempts it as a house hard rule, which is the one exception.
25
+
26
+ ---
27
+
28
+ ## Step 0 -- Hard-rule exemption (check this FIRST)
29
+
30
+ If the FINDING JSON has a `rule` field set to `TS-1`, `TS-2`, `GQL-1`, or
31
+ `PR-1`, do NOT run the five adversarial challenges below. They ask
32
+ runtime-failure questions that a standards violation can never answer, and
33
+ answering them honestly would drop a finding that house-rules policy declares
34
+ Critical on standards grounds rather than on runtime behaviour.
35
+
36
+ Instead run exactly two checks:
37
+
38
+ 1. **Anchor**: the flagged code genuinely appears on a `+` line of
39
+ `DIFF_FILE` (for PR-1, the PR title genuinely lacks a ticket present in
40
+ the commit history). Use the same technique as Step 1:
41
+ ```bash
42
+ grep "^+" "$DIFF_FILE" | grep "<snippet>"
43
+ ```
44
+ 2. **Rule applicability**: the code really does violate the rule as written
45
+ in `references/house-rules.md` -- e.g. an `as const` is not a type cast in
46
+ the TS-1 sense; a `nodes` query that legitimately fetches a single known
47
+ node with a documented comment may satisfy GQL-1; a `.js` file inside a
48
+ Liquid theme repo is exempt from TS-2.
49
+
50
+ Outcome: both checks pass -> `confirmed` at Critical, with `reasoning` naming
51
+ the rule and quoting the violating snippet. Either check fails -> `dropped`,
52
+ with `reasoning` naming which check failed. There is no `downgraded` outcome
53
+ for a hard rule: it either violates the rule or it does not.
54
+
55
+ Then **stop** -- do not continue to Step 1 or the five challenges.
56
+
57
+ ---
58
+
59
+ ## Step 1 -- Diff-anchor check
60
+
61
+ Re-read the actual code at `file:line` in `WORKTREE_PATH` (full context,
62
+ approximately 20-30 lines around the flagged line). Confirm:
63
+
64
+ - The line genuinely appears as a `+` line in `DIFF_FILE`.
65
+ **If the bad code only appears on a `-` line (being removed), the PR is
66
+ already fixing it -- this is not a finding. Drop immediately.**
67
+ - The issue is not already mitigated by surrounding code the review agent did
68
+ not have in context.
69
+
70
+ To check whether it is a `+` line:
71
+
72
+ ```bash
73
+ grep "^+" "$DIFF_FILE" | grep "<flagged_symbol_or_snippet>"
74
+ ```
75
+
76
+ If the flagged code is absent from `+` lines: **drop**.
77
+
78
+ ---
79
+
80
+ ## Step 2 -- Rename / missing-update false positive check
81
+
82
+ When the finding claims "symbol X was renamed but consumer Y was not updated,"
83
+ you MUST run:
84
+
85
+ ```bash
86
+ grep "^+" "$DIFF_FILE" | grep "<new_symbol_name>"
87
+ ```
88
+
89
+ If the replacement already appears in `+` lines of the diff (i.e., the consumer
90
+ was updated in the same PR), this is a **false positive** -- drop silently.
91
+ Only confirm if the new symbol is absent from all `+` lines in consumer files.
92
+
93
+ ---
94
+
95
+ ## Step 3 -- Internal library source check
96
+
97
+ If the finding depends on how a `your org's shared internal` package behaves (message
98
+ format, error shape, method signature), do NOT infer from indirect evidence
99
+ (Sentry titles, email subjects, other projects). Before confirming, run:
100
+
101
+ ```bash
102
+ find <your local workspace root, if any> -name "*.ts" ! -path "*/node_modules/*" \
103
+ | xargs grep -l "<ClassName or symbol>" 2>/dev/null | head -5
104
+ ```
105
+
106
+ Read the actual implementation. If you cannot find the source and cannot prove
107
+ the claim from the code in `WORKTREE_PATH`, **drop the finding** -- "I could not
108
+ find the source" is not a finding.
109
+
110
+ ---
111
+
112
+ ## Step 4 -- Convention findings: precedent check
113
+
114
+ For every finding from the conventions dimension, the cited precedent MUST be
115
+ real. Re-open the `file:line` the finding cites as the existing pattern (in
116
+ `WORKTREE_PATH` or a sibling repo, if you keep one checked out locally) and confirm
117
+ the idiom actually lives there. If the helper/type/convention does not exist
118
+ where claimed, the finding is a taste suggestion in disguise -- **drop it.** A
119
+ convention finding with no verifiable precedent does not ship.
120
+
121
+ ---
122
+
123
+ ## Step 5 -- Fix validation
124
+
125
+ Before confirming the finding, identify the callers and consumers of the changed
126
+ code in `WORKTREE_PATH`. Ask: *"Does the proposed fix break anything
127
+ downstream?"* A fix that silently removes a designed feature (e.g., turning a
128
+ synchronous result into fire-and-forget when the caller renders the result) is
129
+ worse than the original bug. If the fix requires changes beyond the single file,
130
+ note that in `reasoning` rather than silently showing an incomplete patch -- but
131
+ this does not by itself cause a drop.
132
+
133
+ ---
134
+
135
+ ## Step 6 -- Severity test for Critical
136
+
137
+ A finding is Critical only if the failure path is reachable under **normal
138
+ operating conditions**, not only in a worst-case or adversarial scenario. Ask:
139
+ *"Would this fail on a typical production execution today?"* If the answer is
140
+ *only under specific conditions* (large dataset, concurrent load, wrong env
141
+ config), that is **Important**. Reserve Critical for failures that are
142
+ unconditional or highly likely given the feature's intended use.
143
+
144
+ ---
145
+
146
+ ## The five adversarial challenges
147
+
148
+ Every candidate finding in scope must survive five explicit challenges before
149
+ it is allowed in its severity tier. Work through every challenge question in
150
+ order. Write the answer to yourself before deciding the outcome. If you cannot
151
+ answer a question with evidence from the diff and the worktree, that is itself a
152
+ signal the claim is uncertain.
153
+
154
+ ---
155
+
156
+ ### Challenge 1 -- Unconditional reachability
157
+
158
+ > "Under normal production use of this feature, does this failure path trigger
159
+ > without any unusual setup, timing, or configuration?"
160
+
161
+ Re-read the exact code path that leads to the failure. Trace it from the entry
162
+ point (route handler, webhook, cron job, etc.) through the changed code.
163
+
164
+ If the answer is "only under concurrent load" -> **Important**
165
+ If the answer is "only with a specific bad input the API already validates" -> **drop**
166
+ If the answer is "only in a specific env config that is not the default" -> **Important**
167
+ If the answer is "yes, on any normal execution of this feature" -> **passes**
168
+
169
+ ---
170
+
171
+ ### Challenge 2 -- Mitigation in wider context
172
+
173
+ > "Is there any guard, validation, transaction, retry, or framework behaviour
174
+ > within 50 lines of the finding that already prevents or contains this
175
+ > failure?"
176
+
177
+ Re-read **50 lines** around the flagged line in `WORKTREE_PATH` (not just 20-30).
178
+ Also check the caller one level up. Agents often miss:
179
+ - Wrapping `try/catch` in the caller
180
+ - Prisma implicit transactions on `$transaction` calls
181
+ - Zod validation at the route boundary that prevents the bad value from
182
+ ever reaching this code
183
+ - A feature flag / kill switch that keeps this path inactive at deploy time
184
+
185
+ If mitigation exists and fully covers the failure -> **drop** (finding already
186
+ handled)
187
+ If mitigation exists but only partially -> **downgrade to Important** + note the
188
+ gap in `reasoning`
189
+ If no mitigation exists -> **passes**
190
+
191
+ ---
192
+
193
+ ### Challenge 3 -- PR direction
194
+
195
+ > "Is this PR improving this situation compared to before, even if not
196
+ > fully fixing it?"
197
+
198
+ Run:
199
+ ```bash
200
+ grep "^-" "$DIFF_FILE" | grep "<key symbol from finding>"
201
+ ```
202
+
203
+ If the `-` lines show the PR is removing or reducing the problem, this is a
204
+ step forward -- the code was already broken before this PR. Reporting it as
205
+ Critical on a PR that is making things better is misleading.
206
+
207
+ If the PR introduced the problem fresh -> **passes**
208
+ If the PR is reducing an existing problem but not eliminating it -> **Important**
209
+ If the PR is not changing this code path at all (finding traces to unchanged
210
+ code) -> **drop** (not diff-anchored -- should have been caught in the diff-anchor
211
+ check above)
212
+
213
+ ---
214
+
215
+ ### Challenge 4 -- Demonstrability
216
+
217
+ > "Can I write a failing test case in my head -- with specific inputs and
218
+ > expected vs. actual output -- that demonstrates this failure?"
219
+
220
+ This is the strongest signal a Critical finding is real. If you cannot
221
+ articulate the test, you do not fully understand the failure.
222
+
223
+ Write it silently: *"Given input X, function Y returns Z instead of W because
224
+ of line L."*
225
+
226
+ If you can articulate it precisely -> **passes**
227
+ If you can articulate the symptom but not the exact failure mechanism -> **downgrade to Important**
228
+ If you cannot articulate it at all -> **drop** -- "this feels wrong" is not a
229
+ Critical finding
230
+
231
+ ---
232
+
233
+ ### Challenge 5 -- Certainty gate
234
+
235
+ > "Am I relying on any inference, assumption, or indirect evidence that I
236
+ > have not verified from the actual source?"
237
+
238
+ Common failure modes here:
239
+ - Assuming how a `your org's shared internal` package behaves without reading the source
240
+ - Assuming the DB schema based on the model name rather than the Prisma schema
241
+ - Assuming the Shopify API behaviour from memory rather than the docs
242
+ - Assuming a variable is always truthy/falsy without checking where it is set
243
+
244
+ If any assumption is unverified -> resolve it now (read the source, check the
245
+ schema, look up the API). Then re-evaluate.
246
+
247
+ If the finding is fully grounded in code you have read -> **passes**
248
+ If any link in the chain is inference -> **Important at best**, or **drop** if
249
+ the inference is the crux of the claim
250
+
251
+ ---
252
+
253
+ ## Outcome rules
254
+
255
+ The one exception to "a Critical must pass ALL FIVE challenges" is the
256
+ hard-rule path in Step 0: a `rule`-tagged finding is judged solely on the
257
+ anchor + rule-applicability checks and never runs the five challenges.
258
+
259
+ A Critical finding must pass ALL FIVE challenges. Any failure downgrades:
260
+ - One downgrade -> **Important**
261
+ - Two or more downgrades, or any **drop** verdict -> remove from Critical entirely
262
+ (either move to Important with the reason noted, or drop)
263
+
264
+ If the finding survives all five: keep as Critical. Write the verdict
265
+ in one sentence: *"This is Critical because it fails unconditionally on any
266
+ [webhook/sync/button press] due to [specific mechanism], with no mitigation, and
267
+ the PR introduced it."*
268
+
269
+ If you cannot write that sentence cleanly -> it is not Critical.
270
+
271
+ Apply the same five challenges to Important findings when there is doubt --
272
+ the downgrade path is "Important -> Observation -> drop" for that tier.
273
+
274
+ ---
275
+
276
+ ## Output mapping
277
+
278
+ Return EXACTLY one JSON object matching VERDICT_SCHEMA:
279
+
280
+ ```json
281
+ {
282
+ "verdict": "confirmed" | "downgraded" | "dropped",
283
+ "newSeverity": "important" | "observation", // only when verdict is "downgraded"
284
+ "reasoning": "<one sentence stating outcome and why>"
285
+ }
286
+ ```
287
+
288
+ - **confirmed**: finding passed all applicable checks and all five challenges.
289
+ `reasoning` = the one-sentence Critical verdict (or equivalent for Important).
290
+ - **downgraded**: finding survived but at a lower severity due to one downgrade
291
+ outcome above. Set `newSeverity` to the resulting tier. `reasoning` = which
292
+ challenge caused the downgrade and why.
293
+ - **dropped**: finding failed any check -- bad diff anchor, false positive,
294
+ unverifiable precedent, two or more downgrade outcomes, or could not articulate
295
+ the failure. `reasoning` = one sentence naming the disqualifying check.
@@ -0,0 +1,143 @@
1
+ # artifact-page.md -- lekker-review -- living web artifact procedure
2
+
3
+ Execute this procedure AFTER the review file has been saved to `~/code-reviews/`,
4
+ as part of the MAIN loop (not a reviewer/verifier agent). Skip the entire
5
+ procedure silently when `--no-artifact` was passed.
6
+
7
+ This publishes the review as a claude.ai Artifact page whose URL stays STABLE
8
+ across re-reviews of the same PR -- the author watches findings flip from open
9
+ to fixed, commit after commit, at one link. Rendering is delegated to a single
10
+ background subagent so no session-model (Fable) tokens are spent writing HTML.
11
+
12
+ ---
13
+
14
+ ## Step 1 -- Gather inputs
15
+
16
+ All of the following are already in hand after Step 3 of SKILL.md:
17
+
18
+ - `findings.json` path (scratchpad) -- each finding carries: `file`, `line`,
19
+ `severity`, `title`, `description`, `badCode`, `fix`, `rule?`, `precedent?`,
20
+ `agreedBy?`, `verifierReasoning?`, `proof?` (proof = `{attempted, proven,
21
+ reason, testCode?, testCommand?, redOutput?}`).
22
+ - The saved review file path: `~/code-reviews/YYYY-MM-DD-pr-N-repo.md`.
23
+ - PR metadata: `REPO_SLUG`, `PR_NUMBER`, `PR_URL`, title, author, `headRefName`
24
+ → `baseRefName`, head sha, depth, verdict, `isDraft`, `mergeStateStatus`, CI
25
+ status.
26
+ - `PREV_ARTIFACT_URL` -- `null` on first review; on re-review, extracted from
27
+ the prior review file's `**Artifact:** <url>` header line (SKILL.md Step 0
28
+ handles the extraction).
29
+ - Since-last-review data when in re-review mode (fixed vs. still-open lists).
30
+ - The `--no-artifact` flag -- when set, skip this whole procedure silently.
31
+
32
+ ---
33
+
34
+ ## Step 2 -- Launch the renderer subagent (background)
35
+
36
+ Launch exactly ONE `general-purpose` agent, `model: sonnet`, in the
37
+ background. Do not block printing the review on it (see Hard rules).
38
+
39
+ Brief template -- copy verbatim, filling in `<placeholders>`:
40
+
41
+ ```
42
+ Load the `artifact-design` skill (Skill tool) FIRST -- mandatory before
43
+ writing any HTML.
44
+
45
+ Read the findings at <findings.json path> and the saved review at
46
+ <review file path>.
47
+
48
+ Write ONE self-contained HTML file to the session scratchpad:
49
+ review-pr-<PR_NUMBER>-<repo-short>.html
50
+
51
+ Follow the page specification below exactly. Then publish it with the
52
+ Artifact tool using:
53
+ - favicon: "🥩" (keep this IDENTICAL on every republish)
54
+ - title: "Review: <repo-short> #<PR_NUMBER>"
55
+ - description: <one sentence, e.g. "Living code review for PR #<N> --
56
+ findings update as commits land">
57
+ - url: <PREV_ARTIFACT_URL> (include ONLY when non-null, so the SAME
58
+ artifact updates in place instead of minting a new URL; omit the `url`
59
+ parameter entirely on first publish)
60
+
61
+ Return ONLY the resulting artifact URL as your final text. No other output.
62
+
63
+ --- PAGE SPECIFICATION ---
64
+ <paste Step 3 verbatim here>
65
+ ```
66
+
67
+ ---
68
+
69
+ ## Step 3 -- Page specification (the subagent's design contract)
70
+
71
+ Hard requirements for the HTML page:
72
+
73
+ - **Self-contained and theme-aware.** No external fonts/scripts/images/CDNs
74
+ (CSP blocks them). Honor `prefers-color-scheme` as the default signal AND
75
+ `:root[data-theme="dark"]` / `:root[data-theme="light"]` overrides.
76
+ - **Header card:** verdict badge (✅ LGTM / ⚠️ LGTM with changes / 🚫 Needs
77
+ work -- color-coded), PR title linking to `PR_URL`, author, branch → base,
78
+ head sha (short, `code` style), depth chip, findings count by severity, CI
79
+ status. When `isDraft`: a DRAFT banner. When the PR title's ⛔ CANNOT MERGE
80
+ block applies (see output-format.md): an unmissable banner above everything
81
+ else on the page.
82
+ - **Timeline section (re-review mode):** "Since last review" -- one row per
83
+ prior finding: ✅ fixed (title struck through) or ⚠️ still open, each with
84
+ `file:line`. This is the living part of the page. On first review, show
85
+ "First review of this PR" with the head sha and the review's date filled in
86
+ from the review file (do not leave the date as a literal placeholder in the
87
+ output).
88
+ - **Findings**, grouped Critical → Important → Observations → Idiomatic. Each
89
+ finding is an expandable `<details>` block:
90
+ - summary row: severity dot + title + `file:line`
91
+ - body: description, `badCode` in a `<pre>`, `fix` in a `<pre>`, a small
92
+ hard-rule chip (e.g. "TS-1 · hard rule") when `rule` is set, `agreedBy`
93
+ chips when 2+ dimensions agreed, `verifierReasoning` as a muted footnote.
94
+ - **Proof block:** when `proof.proven === true`, render a visually distinct
95
+ "PROVEN -- failing test ran in the worktree" panel: `redOutput` in a
96
+ `<pre>`, a one-line explanation that the test asserts correct behavior, and
97
+ `testCode` collapsed behind its own `<details>`. This is the page's
98
+ centerpiece -- make it prominent (e.g. a red left border) but not garish.
99
+ - **Test Quality + Review Cost sections**, mirrored from the review file,
100
+ kept concise (verdict + gaps + cost table; no need to reproduce every
101
+ sentence).
102
+ - **Layout:** every code block (`badCode`, `fix`, `redOutput`, `testCode`)
103
+ horizontally scrollable within its own container; the page body itself must
104
+ never scroll horizontally. Keep total page weight sane -- plain `<pre>`
105
+ with a theme-aware background, no syntax-highlighting library.
106
+ - **No praise language anywhere** (house rule, same as the saved review):
107
+ state what the code does and the risks it carries, never compliment it.
108
+
109
+ ---
110
+
111
+ ## Step 4 -- Record the URL (fresh post-condition)
112
+
113
+ When the subagent returns:
114
+
115
+ **On success (URL returned):**
116
+ 1. Append or refresh a `**Artifact:** <url>` line in the saved review file's
117
+ header block. Idempotent: replace an existing `**Artifact:**` line if
118
+ present, never duplicate it.
119
+ 2. Re-read the review file to confirm the line landed -- this is the receipt
120
+ (per VERIFICATION.md: a write is not done until re-read from source).
121
+ 3. Print one line to the user: `🔗 Living review: <url>`
122
+
123
+ **On failure (no URL, or the agent errored):**
124
+ - Say so in one line to the user.
125
+ - Do NOT fail or roll back the review over this -- the saved review file is
126
+ the source of truth; the artifact is an enhancement layered on top of it.
127
+
128
+ ---
129
+
130
+ ## Hard rules
131
+
132
+ - The renderer subagent runs in the BACKGROUND -- never block printing or
133
+ delivering the review on its completion.
134
+ - Same PR → same URL, always. Pass `PREV_ARTIFACT_URL` whenever one exists. A
135
+ fresh URL minted for a PR that already has one is a bug -- treat it as such
136
+ if you spot it in the receipt.
137
+ - The artifact is private by default; sharing it is the user's decision, not
138
+ the skill's.
139
+ - Never include secrets or tokens from context. The page contains only what
140
+ the saved review file already contains -- nothing pulled fresh from
141
+ scratchpad context that isn't already in the review.
142
+ - Favicon stays `🥩` forever for this artifact -- never change it on
143
+ republish.