@olegkoval/agent-skills 1.43.0 → 1.44.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (49) hide show
  1. package/adapters/claude/olko-github-pr/skills/lekker-review/SKILL.md +85 -2
  2. package/adapters/claude/olko-github-pr/skills/lekker-review/references/pricing.json +9 -0
  3. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-conventions.md +158 -0
  4. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-implementation.md +76 -0
  5. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-quality.md +87 -0
  6. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-simplification.md +63 -0
  7. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-test-quality.md +172 -0
  8. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/phase4-comparison.md +42 -0
  9. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/profile.md +347 -0
  10. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/profiles/lekker-deep.md +82 -0
  11. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/profiles/lekker-medium.md +81 -0
  12. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/fixtures/context.json +3 -0
  13. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/fixtures/house-rules.md +11 -0
  14. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/fixtures/revmux-report.json +179 -0
  15. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/install-revmux-prompts.sh +66 -0
  16. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/revmux-adapter.mjs +261 -0
  17. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/revmux-engine.sh +156 -0
  18. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/selftest.mjs +76 -0
  19. package/package.json +1 -1
  20. package/plugins/olko-apple-kit/.claude-plugin/plugin.json +1 -1
  21. package/plugins/olko-creative/.claude-plugin/plugin.json +1 -1
  22. package/plugins/olko-garmin-kit/.claude-plugin/plugin.json +1 -1
  23. package/plugins/olko-git-tools/.claude-plugin/plugin.json +1 -1
  24. package/plugins/olko-github-pr/.claude-plugin/plugin.json +1 -1
  25. package/plugins/olko-github-pr/skills/lekker-review/SKILL.md +85 -2
  26. package/plugins/olko-github-pr/skills/lekker-review/references/pricing.json +9 -0
  27. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-conventions.md +158 -0
  28. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-implementation.md +76 -0
  29. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-quality.md +87 -0
  30. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-simplification.md +63 -0
  31. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-test-quality.md +172 -0
  32. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/phase4-comparison.md +42 -0
  33. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/profile.md +347 -0
  34. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/profiles/lekker-deep.md +82 -0
  35. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/profiles/lekker-medium.md +81 -0
  36. package/plugins/olko-github-pr/skills/lekker-review/scripts/fixtures/context.json +3 -0
  37. package/plugins/olko-github-pr/skills/lekker-review/scripts/fixtures/house-rules.md +11 -0
  38. package/plugins/olko-github-pr/skills/lekker-review/scripts/fixtures/revmux-report.json +179 -0
  39. package/plugins/olko-github-pr/skills/lekker-review/scripts/install-revmux-prompts.sh +66 -0
  40. package/plugins/olko-github-pr/skills/lekker-review/scripts/revmux-adapter.mjs +261 -0
  41. package/plugins/olko-github-pr/skills/lekker-review/scripts/revmux-engine.sh +156 -0
  42. package/plugins/olko-github-pr/skills/lekker-review/scripts/selftest.mjs +76 -0
  43. package/plugins/olko-github-pr/skills/lekker-review/workflow.js +43 -6
  44. package/plugins/olko-obsidian/.claude-plugin/plugin.json +1 -1
  45. package/plugins/olko-product/.claude-plugin/plugin.json +1 -1
  46. package/plugins/olko-reflection/.claude-plugin/plugin.json +1 -1
  47. package/plugins/olko-release/.claude-plugin/plugin.json +1 -1
  48. package/plugins/olko-skill-meta/.claude-plugin/plugin.json +1 -1
  49. package/plugins/olko-web-ops/.claude-plugin/plugin.json +1 -1
@@ -0,0 +1,172 @@
1
+ ---
2
+ description: whether tests actually catch bugs — mutation-slip analysis, mock smell, Teifi test conventions
3
+ ---
4
+ ## Lens: lekker-test-quality
5
+
6
+ Review the change with a strict focus on test quality. This is NOT about
7
+ coverage numbers — it is about whether the tests actually catch bugs.
8
+
9
+ Step 1 — Inventory the tests. Read the diff at `{{SCOPE}}`. List every test
10
+ file/spec added or modified. If no test files are in the diff, note that and
11
+ continue to axis (f) below.
12
+
13
+ Step 2 — For each changed test file, read the full file from `{{WORKDIR}}`.
14
+
15
+ Step 3 — Evaluate each of these axes:
16
+
17
+ a) Meaningful assertions vs. smoke tests
18
+ - Does the test verify a specific outcome, or just that no exception was
19
+ thrown / a function returned truthy?
20
+ - Are assertions on the correct thing? (e.g. checking the return value vs
21
+ a side-effect that is the actual goal of the operation)
22
+ - Are there tautological assertions that are always true regardless of the
23
+ implementation? (e.g. `expect(true).toBe(true)`,
24
+ `expect(x).toBeDefined()` when x is always defined by construction)
25
+
26
+ b) Condition coverage
27
+ - Is the happy path tested?
28
+ - Are failure / error paths tested? (invalid input, external API error,
29
+ empty collection, null/undefined, 0 or negative numbers, pagination edge
30
+ cases, missing env vars)
31
+ - For every new if/else or switch in the changed business logic, is each
32
+ branch exercised by at least one test case?
33
+
34
+ c) Regression tests
35
+ - If the change fixes a bug (the ticket mentions a bug, or the diff
36
+ contains "fix" language), is there a regression test that would have
37
+ caught the original bug? Flag if absent.
38
+ - If the change adds a new feature, are the known edge cases of that feature
39
+ tested?
40
+
41
+ d) Mutation-slip analysis (mental mutation testing)
42
+ For the most critical assertions in the test suite, ask: would a simple
43
+ mutation in the production code slip through undetected?
44
+
45
+ Consider these mutation classes:
46
+ - Off-by-one: `> N` changed to `>= N`
47
+ - Operator flip: `&&` to `||`, `===` to `!==`
48
+ - Missing null/undefined guard: remove a `?? default`
49
+ - Wrong variable: using `a` where `b` was intended
50
+ - Return-value swap: returning the wrong field from an object
51
+ - Early-return removed: a guard clause deleted
52
+
53
+ For each mutation class relevant to the changed business logic, determine
54
+ whether at least one test assertion would catch it. Summarize as a short
55
+ paragraph: "Mutations that would slip through: ..." or "No obvious
56
+ mutation-slip gaps found."
57
+
58
+ e) Test isolation and reliability
59
+ - Do tests share mutable state across cases without resetting between runs
60
+ (a beforeEach that does not clean up)?
61
+ - Are there tests that depend on execution order or global singletons?
62
+ - Could a test make a real network/DB call in CI (flaky)? The fix is a
63
+ simple injected fake at the boundary, not blanket module mocking — see
64
+ axis (g).
65
+ - Are async tests properly awaited? (floating promises, missing `await` on
66
+ `expect().resolves`, unhandled rejections)
67
+
68
+ f) Test-to-code ratio signal
69
+ If the diff adds > 50 lines of new business logic with zero new or modified
70
+ test files, flag it explicitly. Then check whether existing test files
71
+ already cover the new code paths:
72
+ ```bash
73
+ find <workdir> -name "*.test.ts" -o -name "*.spec.ts" | \
74
+ xargs grep -l "<key changed symbol>" 2>/dev/null | head -5
75
+ ```
76
+ Report whether existing coverage closes the gap or not.
77
+
78
+ g) Mock smell — test the behavior, not the way it's built
79
+ (House standard: Notion "To mock or not to mock" —
80
+ https://app.notion.com/p/teifi/To-mock-or-not-to-mock-36ff8ed0f7db80f09495d174e3b86cd6)
81
+ Flag tests coupled to *how* the code is built rather than *what* it does for
82
+ the user. For each smell, do NOT just criticize: give the concrete no-mock
83
+ refactor. The default fix is almost always the same shape — pull the logic
84
+ into a pure function (functional core) and test that directly, leaving a thin
85
+ shell covered by a few real integration tests.
86
+
87
+ - Mocking IO just to reach logic: a GraphQL/fetch/DB client mocked only so a
88
+ test can read a computed value back. It couples to the query shape AND the
89
+ markup while barely testing the logic, and silently rots as the real
90
+ dependency drifts from the mock. This applies just as much on the backend
91
+ as in a component: mocking a service client (e.g. `vi.mocked(epicorClient)`
92
+ returning a canned async generator/array) just to reach a reassembly loop
93
+ or a multi-stream join is the same smell as mocking `fetch` in a component.
94
+ → Extract the computation into a pure function over plain data and test
95
+ that with literals (no mocks, no render, no mocked client). Cover
96
+ fetch-and-wire once with a real integration test. For streaming/paging
97
+ code specifically: a shared `collect()`/reassembly helper should be a
98
+ pure function of `AsyncGenerator<T[]> → Promise<T[]>` (or similar) fed a
99
+ plain fake generator in its own test — never a mocked client — and any
100
+ call-site logic that combines multiple streams (parallel joins, chunked
101
+ batching) should be its own pure function tested the same way.
102
+ - Spying on calls / asserting call shape: `toHaveBeenCalledWith`,
103
+ `toHaveBeenCalledTimes`, a `vi.fn()` used as a probe. Asserts *how* a
104
+ function was called, freezing batching/page-size/call-count in place even
105
+ when the output is identical.
106
+ → Assert the output, never the call log. If a boundary is genuinely
107
+ needed, inject a simple fake (a plain function returning canned data),
108
+ not a spy.
109
+ - Mocking a component to read its props back (re-emitting props as `data-*`
110
+ attributes, then asserting on them): the assertions are about the mock, and
111
+ break on a component swap or prop rename that changes nothing a user sees.
112
+ → Pull the logic (pagination, display state) into a pure function over
113
+ plain values — e.g. `paginate(items, page, pageSize)` — and assert the
114
+ value it returns.
115
+ - Mocking a query hook (`useQuery` / a `use-X` hook) to hand a component
116
+ canned data: re-tests React Query's own plumbing and couples to the hook's
117
+ return shape.
118
+ → Keep the `queryFn` as IO (integration-tested); move the transform into a
119
+ pure function wired through React Query's `select` option, and unit-test
120
+ that pure function with plain data.
121
+ - Asserting implementation details: that a component is memoized, that a
122
+ specific child renders, that work happens through a fixed sequence of
123
+ calls.
124
+ → Delete the assertion; assert the user-visible behavior instead.
125
+
126
+ When a mock IS the right call, do NOT flag it: a real IO boundary that must
127
+ be exercised where a fake is impractical, non-determinism that must be pinned
128
+ (time, randomness, injected network failure), or a dependency that genuinely
129
+ cannot run in the test environment. The tell for a good mock: the behavior
130
+ under test only exists because the boundary did something (e.g. a retry
131
+ banner that appears only when the fetch rejects). Even then, prefer an
132
+ injected simple fake over a module-level spy, and assert what the user sees.
133
+
134
+ h) Teifi test conventions
135
+ Read `{{PROFILE}}` and apply its §4 (test conventions) to every test file in
136
+ the diff: `__tests__/` placement beside the source, `.test.ts` vs
137
+ `.test.tsx`, English-sentence test names, `describe` nesting <= 2,
138
+ `afterEach(cleanup)`, `retry: false` on the test QueryClient (major — its
139
+ absence hangs the suite on a failing query), the getByRole → getByLabelText
140
+ → getByText → getByTestId priority, module-boundary mocking, `safeParse` for
141
+ Zod schemas, and the typed `it.each` matrix with AC tags plus the
142
+ exhaustiveness assertion whenever behaviour depends on 3+ independent inputs.
143
+ A `vi.mock('@lib/common/…')` inside `common/` is major, not style: the
144
+ alias resolves only from `web/`, so the mock silently does nothing and the
145
+ test passes against the real module. Quote the line, give the relative path
146
+ as the fix.
147
+ A comment explaining WHY a test exists (a captured past bug, a subtle
148
+ contract) is sanctioned — never flag it as over-commenting.
149
+
150
+ Report each per-line issue as:
151
+ `test-file:line — <concise description of the gap or weakness>`
152
+
153
+ For every mock-smell finding from axis (g), append the no-mock fix on the next
154
+ line as `→ Fix: <pure-function / simple-fake refactor in one line>`. A
155
+ criticism without a fix is incomplete.
156
+
157
+ Report the mutation-slip analysis as a single paragraph under a
158
+ "**Mutation-slip risk:**" heading — not as line items.
159
+
160
+ Rules:
161
+ - Every per-line finding must trace to test code in the diff, OR to business
162
+ logic added in the diff that has no test coverage at all.
163
+ - Report problems only. No praise for tests that meet the bar.
164
+ - If no test files are changed AND no existing tests cover the new code paths,
165
+ report: "No test coverage for new code paths."
166
+ - Quote the verbatim offending line(s) — never paraphrased, never reconstructed
167
+ from memory. When the finding IS the absence of a test, there is no test line
168
+ to quote: quote the untested production line(s) from the diff instead, and
169
+ give the test that should exist as the fix. Never drop a "no coverage"
170
+ finding just because nothing bad is written down — absence is the finding.
171
+ - A finding you cannot quote and cannot fix (outside the no-coverage case
172
+ above) is a finding you have not proven — drop it instead.
@@ -0,0 +1,42 @@
1
+ # Phase 4: revmux engine on real PRs
2
+
3
+ One row per live `--engine revmux` run. Numbers come from revmux `stats` via
4
+ `scripts/revmux-adapter.mjs`; USD is the `pricing.json` list-price placeholder
5
+ (`verified: false`, output price applied to every token, synth+verify unpriced),
6
+ and every run so far went through the Max subscription, so USD is a relative
7
+ signal only. "Confirmed" means the main loop kept the finding after reading the
8
+ worktree; "merged" means two revmux findings described one mechanism.
9
+
10
+ | Date | PR | Depth | Wall | Tokens | USD est. | revmux findings | Kept | Re-severity | Hard-rule re-promotions | Merged | False positives |
11
+ |---|---|---|---|---|---|---|---|---|---|---|---|
12
+ | 2026-09-09 | evi-integrations #544 | medium (auto) | 15m40s | 12,604,506 | $169.53 | 4 | 3 (1 C / 2 I) + obs | 1 important → critical (lost role grants, watermark written before post-pass) | 0 | 0 | 0 |
13
+ | 2026-09-09 | evi-integrations #537 | deep (auto, *.sql) | 11m08s | 11,245,064 | $141.61 | 4 | 3 (0 C / 2 I / 1 idiomatic) + obs | none | 0 | 1 (#1 reactivation + #4 guard removal → one Important) | 0 |
14
+
15
+ ## Notes per run
16
+
17
+ ### #544
18
+
19
+ - revmux under-rated the watermark ordering bug as important; the main loop
20
+ raised it to Critical after tracing `executeSyncStrategy` writing the CUSTOMER
21
+ watermark before the role post-pass.
22
+ - One finding referenced `getShopFlag`, which does not exist in this repo;
23
+ rewritten to an env toggle with Reflag as the policy target. Class: fix code
24
+ invented from another repo's helper. Worth a lens note.
25
+ - Depth medium, so verify ran on Criticals only; revmux's own verify stage
26
+ covered all four.
27
+
28
+ ### #537
29
+
30
+ - Two of four findings were the same mechanism seen from two lenses (bugs,
31
+ adversarial). Synthesis did not merge them; the main loop did.
32
+ - All four confirmed against the worktree. The idiomatic one (comment
33
+ narrating deleted code) is exactly the teifi-conventions §2 class.
34
+ - Reactivation and soft-delete files are outside the diff, so the main
35
+ finding could only be anchored on the sync-path signal it replaces.
36
+
37
+ ## Still open
38
+
39
+ - Same PRs through the default `workflow` engine for a side-by-side (Oleg-driven).
40
+ - Read-only enforcement: revmux default `--tools` includes Bash; prompt-enforced
41
+ only. `--tools=Read,Grep,Glob,WebFetch,WebSearch` override not yet applied.
42
+ - Adapter: price synth + verify once revmux reports a per-model split.
@@ -0,0 +1,347 @@
1
+ # Teifi project profile
2
+
3
+ This is the Teifi project profile, handed to revmux as `{{PROFILE}}` for the lekker-medium
4
+ and lekker-deep review profiles. It concatenates the two files lekker-review itself reads as
5
+ `teifi-rules.md` (hard rules, always critical) and `teifi-conventions.md` (soft house style).
6
+
7
+ ---
8
+
9
+ # Teifi rules, taxonomy, and stack context
10
+
11
+ ## Teifi Hard Rules (apply during review AND development)
12
+
13
+ These rules are non-negotiable. Violations are always **Critical** findings regardless of depth or other filters. Apply them when developing a feature or bugfix, not only during review.
14
+
15
+ ### TS-1 — Type safety (TypeScript only)
16
+
17
+ - No type casting (`as X`, `<X>expr`) — ask them if they are Harry Potter for casting spells.
18
+ - No `any` — except in test files where types are genuinely hard to express; even there, blatantly omitted types (e.g. `any[]` on a known shaped list) must be flagged.
19
+ - Every finding: quote the cast/`any`, explain the correct type, show the fix.
20
+
21
+ ### TS-2 — No JavaScript files
22
+
23
+ - No `.js` files may be added to any Teifi integrations repo.
24
+ - Exception: Liquid themes (Online Store 2.0 Shopify themes) may contain `.js`.
25
+ - If the PR adds a `.js` file to a non-theme repo, flag it as Critical: must be converted to `.ts`.
26
+
27
+ ### GQL-1 — GraphQL NodesConnection pagination
28
+
29
+ - Every query that uses a nodes connection (`nodes { ... }`) **must** include `pageInfo { hasNextPage endCursor }` alongside the nodes.
30
+ - All remaining pages **must** be fetched — a single-page fetch with no loop/recursion is a bug.
31
+ - The page size **must** be `250` (Shopify max). If any other value is used, a code comment explaining why is required; if no comment exists, flag it.
32
+
33
+ ### PR-1 — PR title must be prefixed with Linear ticket(s)
34
+
35
+ - PR title must start with `[GIC-123]` (or the relevant project prefix) in square brackets.
36
+ - Go through the commit history: if merged PRs or commits reference Linear tickets in square brackets (`[GIC-123]`), all of them must appear comma-separated in the current PR title (e.g. `[GIC-123,GIC-124]`).
37
+ - This is a **blocking** finding: display a prominent `⛔ CANNOT MERGE` warning and recommend the correct title prefix. Confidence score is not affected — this is a process rule, not a code quality signal.
38
+
39
+ ### 1g. Repo placement check (Teifi multi-repo projects only)
40
+
41
+ For any PR in a `Teifi-Digital/` repo, verify that the code being changed
42
+ belongs in *this* repo and not a sibling repo.
43
+
44
+ **Teifi repo taxonomy:**
45
+
46
+ | Repo pattern | Purpose |
47
+ |---|---|
48
+ | `*-live` (e.g. `gic-live`, `evi-live`) | Shopify app: customer account extensions, app blocks, storefront extensions, Polaris admin UI |
49
+ | `*-integrations` (e.g. `evi-integrations`; see the GIC exception below) | Backend ERP sync: cron jobs, orchestrator, BC/Sage/Jitterbit/ROI API clients |
50
+ | `teifi-digital` / shared libs | Cross-project utilities, shared types |
51
+
52
+ **Note:** `gic-integrations` has an `extensions/` folder containing legacy/reference extensions (e.g. `link-account-customer`), but **new customer account extensions for GIC should target `gic-live`** per project specs. Always check the Linear/Notion ticket for explicit repo path — don't infer from existing repo contents alone.
53
+
54
+ **Action:** If the diff adds a new Shopify extension and the PR targets `*-integrations`, check the Linear/Notion spec for the explicit target directory. If the spec names `*-live`, flag it as a **Critical** placement error (`## 🏠 Wrong Repo`).
55
+
56
+ Include: spec quote with correct path, which repo to target, and the risk (wrong Shopify Partner app, extension not published to correct store).
57
+
58
+ ## Notes for Teifi / evi-integrations
59
+
60
+ ### Standing coding rules (apply during development and review)
61
+
62
+ | Rule | What | When to flag |
63
+ |---|---|---|
64
+ | TS-1 | No type casts (`as X`), no `any` | Critical; test files lenient on genuine unknowns only |
65
+ | TS-2 | No `.js` files in integrations repos | Critical; Liquid themes exempt |
66
+ | GQL-1 | nodes connections need `pageInfo`, all pages fetched, size=250 | Critical if pageInfo/pagination missing; Important if size≠250 without comment |
67
+ | PR-1 | PR title must start with `[TICKET-NNN]`; include all commit-referenced tickets | Blocking — ⛔ CANNOT MERGE warning |
68
+ | FLAG-1 | Reflag repos only: risky change should ship behind a feature flag | Non-blocking; `important` at most, usually `observation`. Never a `rule:` tag |
69
+
70
+ These apply equally when you are writing a feature or bugfix — not only in review.
71
+
72
+ ### FLAG-1 — Ship behind a feature flag (Reflag repos only)
73
+
74
+ Applies ONLY where a `package.json` (any depth, excluding `node_modules`) depends on
75
+ `@reflag/node-sdk` or `@teifi-digital/reflag-client`. Elsewhere there is no flag client,
76
+ so the finding is unactionable and must not be raised.
77
+
78
+ Flag it when any of these hold:
79
+ - a client should validate it before everyone sees it
80
+ - it changes data shape or what gets written
81
+ - it touches orders, money, or fulfilment
82
+ - it cannot be verified without real client data or volume
83
+
84
+ Just ship it when: pure UI/copy with no data change; a bug fix that is strictly better
85
+ and obviously correct; an internal/admin-only surface; fully covered by tests and
86
+ verifiable in staging.
87
+
88
+ Tie-breaker: would you be comfortable being the one to fix this forward at 2am?
89
+ Yes → ship it. No → flag it.
90
+
91
+ Deliberately NOT blocking, unlike TS-1/GQL-1/PR-1: whether something needs a flag is a
92
+ rollout judgement, not a correctness violation, and a blocking comment on every
93
+ borderline diff trains people to ignore the signal. Never invent a concrete flag key —
94
+ keys must be confirmed against Reflag, so say a flag is needed without naming one.
95
+
96
+ Related: a diff that BOTH adds a column/table AND changes what is read or written must
97
+ be split into expand / migrate / read-switch / contract PRs (`important`, name the split).
98
+
99
+ ### Stack context to inform the review:
100
+
101
+ - **Backend:** TypeScript, Node.js, Express, Prisma, pgtyped, PostgreSQL
102
+ - **Frontend:** React, Shopify Polaris, Vite
103
+ - **Shopify:** REST Admin + GraphQL Admin, Webhooks, Shopify Functions, genql
104
+ - **External APIs:** Business Central (OAuth2, rate-limited), Salesforce GraphQL, ROI
105
+ - **Infra:** Docker Compose locally; environment-var–driven cron syncs via `orchestrator.ts`
106
+ - **Type gen pipeline:** pgtyped (SQL→TS), genql (GraphQL→TS), json2ts (schemas→TS) —
107
+ check that generated files are regenerated when their sources change
108
+ - **MCPs available:** Linear, Slack, Notion, Shopify Dev docs, Harvest, Sentry
109
+
110
+ ### Repo taxonomy (for Step 1g placement check)
111
+
112
+ **Per-project repo structure varies — always verify before flagging.**
113
+
114
+ - For **EVI project**: `evi-integrations` = backend ERP sync only; `evi-live` = Shopify app (extensions, Polaris UI).
115
+ - For **GIC project**: `gic-integrations` owns the ERP backend sync and retains
116
+ legacy/reference Shopify extensions. New GIC customer account extensions belong
117
+ in `gic-live`; do not treat the legacy `extensions/` folder as placement precedent.
118
+ - For other projects (`rsl-*`, `elmt-*`, etc.): check the repo's `extensions/` folder
119
+ presence before assuming a split — do not assume the `*-integrations` pattern always
120
+ means backend-only.
121
+
122
+ **Step 1g action**: Before flagging a placement mismatch, run:
123
+
124
+ ```bash
125
+ gh api "repos/<REPO_SLUG>/git/trees/HEAD" 2>/dev/null | python3 -c "
126
+ import sys,json; t=json.load(sys.stdin).get('tree',[]); print([f['path'] for f in t if f['path']=='extensions'])
127
+ "
128
+ ```
129
+
130
+ If `extensions/` exists in the target repo, placement may be intentional, but
131
+ that alone is not precedent for new GIC extensions. Flag when the repo has no
132
+ `extensions/` folder, or when a GIC spec targets the new extension to `gic-live`.
133
+
134
+ ---
135
+
136
+ # Teifi soft conventions (styling, naming, comments, tests, hygiene)
137
+
138
+ Companion to `teifi-rules.md`. Those are the four HARD rules (TS-1, TS-2,
139
+ GQL-1, PR-1) — always Critical. This file is the house style the Teifi
140
+ plugin skills enforce during development (`teifi-dev:code-review`,
141
+ `rename-pass`, `comment-stripper`, `unit-test-best-practices`). A review that
142
+ misses them lets the same nits come back from the human reviewer.
143
+
144
+ Severity ceiling: everything here is **idiomatic** unless a row says otherwise.
145
+ Cite a precedent from the codebase where the row asks for one, and never
146
+ inflate a style deviation into Critical.
147
+
148
+ ---
149
+
150
+ ## 1. Naming matrix (source: `teifi-dev` rename-pass agent)
151
+
152
+ Flag a name the diff INTRODUCES that breaks a row. Renaming is cheap in the
153
+ diff, expensive later — but a rename is still a cost, so leave conforming
154
+ names alone.
155
+
156
+ ### Values
157
+
158
+ | Kind | Convention | Example |
159
+ |---|---|---|
160
+ | boolean | `is` prefix, no negatives (`has`/`can`/`should` for possession/permission/policy) | `isFrozen` |
161
+ | array | plural noun (value plural, type stays singular) | `variantIds` |
162
+ | keyed lookup | `…ById` suffix | `toneByStatus` |
163
+ | id | `Id` / `Ids` suffix | `shipmentId` |
164
+ | timestamp | `At` suffix | `scheduledAt` |
165
+ | duration | unit suffix | `timeoutMs` |
166
+ | count | `Count` suffix | `lineCount` |
167
+ | casing | camelCase values · PascalCase types/components · `CONSTANT_CASE` constants | |
168
+
169
+ ### Verbs — one per job, chosen by cost + purity
170
+
171
+ `get` cheap/sync · `fetch` async I/O · `create` new persisted entity ·
172
+ `build` assembles in memory (pure) · `derive` pure value from existing state ·
173
+ `parse` raw → typed · `format` value → display string · `update` change a
174
+ persisted entity · `validate` returns validity · `assert` throws · `ensure`
175
+ idempotently makes state hold.
176
+
177
+ A `getX` that does network I/O is a finding (`fetchX`). A `createX` that
178
+ mutates an existing row is a finding (`updateX`).
179
+
180
+ ### Effect affixes — put the surprise in the name
181
+
182
+ `…OrThrow` · `…OrDefault` · `upsert` · `…ForUpdate` (row lock) ·
183
+ `…SkipLocked` · `try…` (returns result, doesn't throw) · `with…`
184
+ (acquire→run→release) · `…Sync` · `…Cached` · `unsafe…`.
185
+
186
+ A function that takes a row lock, throws on miss, or returns a cached value
187
+ without saying so in its name is a finding — that surprise is exactly what the
188
+ next caller will miss.
189
+
190
+ ### React & types
191
+
192
+ - React: `use` hook · `handle` (impl) / `on` (prop) · `Props` type · `with` HOC
193
+ · components PascalCase and **domain-prefixed when collidable**
194
+ (`AdjustmentStatusBadge`, not `StatusBadge`).
195
+ - Types: `Input` / `Output` · `Result` (ok/err) · `Schema` (Zod) ·
196
+ `Brand<string,'X'>` for ids · PascalCase noun, no `I-` prefix, never
197
+ pluralize a type.
198
+
199
+ ### Bare generic nouns
200
+
201
+ `line`, `item`, `node`, `record`, `entry`, `group`, `row`, `value`, `key` are
202
+ findings when the surrounding domain has two or more qualified variants in
203
+ scope (arrival line vs receipt line; source node vs target node). Qualify with
204
+ the domain role — variables, params, fields, type aliases, **type params**
205
+ (`TLine` → `TReceiptLine`), and the functions built on the noun.
206
+
207
+ ### NEVER flag a rename at a boundary
208
+
209
+ DB table/column names (and any `Row`/`Dto` mirroring them), GraphQL / oRPC /
210
+ OpenAPI contract fields, enum string values, route strings, wire/JSON keys.
211
+ Something outside the diff reads them. The app-layer alias may be renamed
212
+ (`createdAt @map("created_at")`); the boundary name may not. A "rename this
213
+ column" finding is a false positive — say the boundary name is bad and leave
214
+ it to a human if it matters.
215
+
216
+ ---
217
+
218
+ ## 2. Comment policy (source: `teifi-dev` comment-stripper + code-review Cat. 5)
219
+
220
+ **Default: a comment should not exist.** It earns its place only by saying
221
+ something the code *cannot*, and then in as few words as possible. Flag each
222
+ offending comment the diff ADDED as its own `idiomatic` finding with the
223
+ deletion as the `fix` — never a vague "too many comments". Pre-existing
224
+ comments are out of scope.
225
+
226
+ Flag a comment that:
227
+
228
+ - restates the next line (`// increment counter` over `counter += 1`);
229
+ - narrates a step or captions a block ("first we fetch, then we map…");
230
+ - **narrates a whole function or type** — JSDoc restating a well-named
231
+ signature, its params, or its return shape. A doc comment earns its place
232
+ only for a non-obvious *contract*;
233
+ - explains obvious syntax or a well-known API;
234
+ - repeats a rationale stated elsewhere in the diff (keep ONE canonical place);
235
+ - is changelog/AI noise — ticket IDs (`EVI-123`), person names, multi-paragraph
236
+ "why we chose X" essays. A single `@see EVI-123` JSDoc tag is fine;
237
+ - **references what the reader cannot see** — a removed line or a prior
238
+ approach ("no longer using the old Y"). Source shows what the code *is*;
239
+ - **documents invisible coupling** — "ordered this way because some other code
240
+ does X". The fix is clearer structure, not a comment enshrining it;
241
+ - **says what a name or type could say** — then the code should carry it (this
242
+ is a naming finding per §1, not a comment to keep).
243
+
244
+ KEEP (do not flag): an external-system quirk, an ordering/concurrency
245
+ constraint, a bug workaround, a footgun, a "looks wrong but is correct
246
+ because…" causal chain, a directive, a license header, and every lint/type
247
+ pragma (`eslint-disable`, `@ts-expect-error`, `prettier-ignore`).
248
+
249
+ Heuristic: a comment about as long as the code it sits on, naming no real
250
+ gotcha, is noise. Code is type-checked; prose is not.
251
+
252
+ ---
253
+
254
+ ## 3. Hygiene & debug artifacts — severity is fixed, do not soften
255
+
256
+ Scan `+` lines only (added by this diff; pre-existing occurrences are out of
257
+ scope).
258
+
259
+ | Artifact | Severity |
260
+ |---|---|
261
+ | `console.log(` / `console.debug(` added to non-CLI production code | important |
262
+ | `debugger;` | critical |
263
+ | `.only` on a test (`it.only`, `describe.only`, `fit(`, `fdescribe(`) — silently skips the rest of the suite | critical |
264
+ | new `TODO:` / `FIXME:` / `HACK:` / `XXX:` with no ticket reference | idiomatic |
265
+ | commented-out code block > 3 lines | idiomatic |
266
+ | hardcoded URL / endpoint that belongs in an env var | important |
267
+ | hardcoded magic number that should be a named constant | idiomatic |
268
+ | unreachable code after `return`/`throw` | important |
269
+
270
+ `console.error`/`console.warn` on a real error path is not a finding unless the
271
+ repo has a logger idiom — then cite it.
272
+
273
+ ---
274
+
275
+ ## 4. Test conventions (source: `teifi-dev:unit-test-best-practices`)
276
+
277
+ These are on top of the mutation/mock analysis the test-quality agent already
278
+ does. Each is an `idiomatic` finding with the corrected code as the `fix`.
279
+
280
+ - **Placement:** tests live in `__tests__/` **next to** the source file.
281
+ `.test.ts` for pure logic, `.test.tsx` when JSX is needed. A test parked in a
282
+ top-level `test/` dir or beside the source without `__tests__/` is a finding.
283
+ - **Names read as English sentences:** `'returns false when name is null'`, not
284
+ `'should return false'`.
285
+ - **`describe` nesting max 2 levels;** flat `it()` blocks are fine for small
286
+ functions. Grouping that adds no clarity is a finding.
287
+ - **`afterEach(cleanup)`** in every RTL test file.
288
+ - **`new QueryClient({ defaultOptions: { queries: { retry: false } } })`** in
289
+ every hook/component test wrapper. Missing `retry: false` silently hangs the
290
+ test on a failing query — flag it as `important`, not idiomatic.
291
+ - **Query priority:** `getByRole` → `getByLabelText` → `getByText` →
292
+ `getByTestId` (last resort only). A `getByTestId` where a role query works is
293
+ a finding.
294
+ - **Mock paths are RELATIVE, never the `@lib/common` alias.** The alias
295
+ resolves only from `web/`; inside `common/` it is silently ignored and the
296
+ mock has no effect — the test then passes against the real module. Flag any
297
+ `vi.mock('@lib/common/…')` inside `common/` as `important`.
298
+ - **Mock at the module boundary,** not `vi.spyOn(mod, '_internal')`.
299
+ - **Zod schemas:** test via `safeParse` and assert `result.success`, not
300
+ try/catch around `parse`.
301
+ - **Matrix tests:** when behaviour depends on 3+ independent boolean/enum
302
+ inputs, drive `it.each` from a typed matrix, tag each row with the AC it
303
+ verifies, and include the exhaustiveness assertion
304
+ (`expect(matrix).toHaveLength(N * M * K)`). A hand-picked subset of a
305
+ combinatorial space is a coverage-gap finding.
306
+ - **Comment the non-obvious scenario:** a test capturing a past bug explains
307
+ *why* (this is a sanctioned comment — never flag it under §2).
308
+
309
+ ---
310
+
311
+ ## 5. Commit & PR hygiene
312
+
313
+ Beyond PR-1 (hard rule). All `idiomatic` — a squash fixes them and they never
314
+ block.
315
+
316
+ - Conventional commits: `<type>(<scope>): <description>` with type in
317
+ `feat|fix|refactor|test|docs|chore|style|perf|build|ci`; description
318
+ lowercase, imperative, no trailing period.
319
+ - Vague subjects (`fix`, `update`, `wip`, `changes`, `stuff`) — flag, recommend
320
+ a rewrite.
321
+ - Many WIP commits — recommend a squash before merge, in one line, once.
322
+ - Never mention Claude Code / the assistant in commit or PR text; a
323
+ `Co-Authored-By: Claude` trailer or "Generated with Claude Code" footer in
324
+ the commit log is a finding.
325
+
326
+ ---
327
+
328
+ ## 6. Generated code — check the source, not the artifact
329
+
330
+ The Teifi type-gen pipeline means several files must move together. When a
331
+ source changes and its generated artifact does not (or vice-versa), that is an
332
+ `important` finding:
333
+
334
+ | Source changed | Artifact that must be regenerated |
335
+ |---|---|
336
+ | `services/db/queries/*.sql` | `services/db/queries/generated/` (pgtyped) |
337
+ | `services/gql/queries/*.graphql` | `services/gql/queries/generated/queries.ts` (genql) |
338
+ | `schemas/*.json` | `schemas/generated/` (json2ts) |
339
+ | `prisma/schema.prisma` | a migration in `prisma/migrations/` + Prisma client |
340
+
341
+ Never review the *content* of a generated file as if it were hand-written — no
342
+ naming, comment, or complexity findings inside `generated/`. Review the source.
343
+
344
+ Related hard convention (project CLAUDE.md): Shopify Admin API calls go through
345
+ the genql client (`gql.<file>.<query>.run(graphql, vars)`) — a hand-rolled
346
+ `fetch` to `/admin/api/…/graphql.json` is an `important` finding even for a
347
+ one-off probe.
@@ -0,0 +1,82 @@
1
+ ---
2
+ description: Teifi deep review — lekker-medium plus revmux's own bugs+impl second opinion, claude-only
3
+ model: claude/sonnet:medium
4
+ agents:
5
+ - {name: quality+impl, lenses: [lekker-quality, lekker-implementation], color: cyan}
6
+ - {name: simpl+conventions, lenses: [lekker-simplification, lekker-conventions], color: magenta}
7
+ - {name: tests, lenses: [lekker-test-quality, tests], color: green}
8
+ - {name: adversarial, lenses: [adversarial], model: claude/sonnet:high, color: yellow}
9
+ - {name: bugs+impl, lenses: [bugs, impl], color: blue}
10
+ stages: {synthesis: claude/opus:medium, verify: claude/sonnet:high}
11
+ ---
12
+ You are one reviewer on a panel. Other reviewers are working the same change in parallel with
13
+ different lenses. You never see their findings and must not guess at them — report what your own
14
+ lenses find.
15
+
16
+ This review is **read-only**. You may read files and run read-only commands such as `git diff`,
17
+ `git log` and `rg`. Do not modify, delete, move, stage or commit anything, and do not write a file
18
+ through a shell redirect. Report what you find; changing it is the caller's job, never yours.
19
+ Do not run tests, builds or the linter - all of that was done before the review and passed.
20
+
21
+ ## Where the context lives
22
+
23
+ Every item below is a **path**, not the text it names. Read the file or directory before you start.
24
+
25
+ - `{{SCOPE}}` — what is under review and the command that produces the diff. Read this first and run
26
+ that command yourself.
27
+ - `{{GOAL}}` — what the change is trying to achieve.
28
+ - `{{PROFILE}}` — Teifi's own rules and conventions. Where they disagree with your general taste,
29
+ they win. This is also where the hard-rule text (TS-1, TS-2, GQL-1, PR-1) and the Teifi
30
+ conventions (naming matrix, comment policy, hygiene severities, test conventions) live in full.
31
+ - `{{CONTEXT}}` — a directory of supporting material: ticket text, design notes, spec excerpts, CI
32
+ status, Sentry signals, existing review comments.
33
+ - `{{WORKDIR}}` — run every command from here.
34
+
35
+ Any of these may read `none provided`. That is not an error and not something to work around: the
36
+ caller supplied nothing for it, so calibrate severity generically to that extent rather than
37
+ inventing the missing context.
38
+
39
+ ## Severity bar
40
+
41
+ No nitpicking. Critical and major findings are reserved for things that could cause bugs, outages,
42
+ data loss, security incidents, or real performance problems at scale.
43
+
44
+ - **critical** — a bug, an outage, data loss, a security hole, or a real performance problem at scale.
45
+ - **major** — wrong behavior, or a broken contract a caller executes against.
46
+ - **minor** — a real, contained defect.
47
+
48
+ Style preference and taste alone are never a finding. Anything you cannot place on that bar is not
49
+ a finding — leave it out.
50
+
51
+ ## Hard-rule findings are policy, not a runtime question
52
+
53
+ Findings titled `[TS-1]`, `[TS-2]`, `[GQL-1]`, or `[PR-1]` are Teifi's own policy violations,
54
+ defined in full in `{{PROFILE}}`. Confirm one when the quoted code shows the pattern the rule
55
+ names — a cast, an `any`, a `.js` file outside a theme repo, a missing `pageInfo`/pagination, a PR
56
+ title missing its ticket prefix. Never rate a hard-rule finding by its runtime impact and never mark
57
+ it immaterial for lack of one: the rule itself is the standard, and violating it is always critical,
58
+ independent of whether it happens to fail at runtime today.
59
+
60
+ ## Reporting
61
+
62
+ Apply every lens you carry, in full, and tag each finding with the lens that raised it.
63
+
64
+ - Point at a specific file and line. A finding with no location cannot be verified.
65
+ - State the failure concretely: the input or state, and what goes wrong because of it.
66
+ - Report the confidence you actually have, not the confidence that keeps the finding alive.
67
+ - Say when a problem is pre-existing rather than introduced by the change under review.
68
+ - Do not report one problem twice under two lenses. Report it once and name both lenses on it.
69
+
70
+ ## What not to report
71
+
72
+ Silence beats a finding the reader has to disprove. Do not report:
73
+
74
+ - a defect on a line this change did not touch, unless the change is what makes it reachable
75
+ - anything a linter, compiler or type checker catches. All of them ran before the review and passed
76
+ - a lint or vet rule the code silences deliberately, with the directive visible
77
+ - a missing test, missing doc or general-quality observation the project's own rules do not ask for
78
+ - a nitpick a senior engineer reading this diff would not raise
79
+ - a behaviour change that is plainly the point of the change
80
+
81
+ Pre-existing problems are the one exception: report them, and say so, so the reader can weigh them
82
+ separately from what the change introduced.