@olegkoval/agent-skills 1.43.1 → 1.45.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (56) hide show
  1. package/adapters/claude/olko-github-pr/skills/lekker-review/SKILL.md +113 -4
  2. package/adapters/claude/olko-github-pr/skills/lekker-review/references/agents/fix-verifier.md +56 -2
  3. package/adapters/claude/olko-github-pr/skills/lekker-review/references/fix-mode.md +19 -1
  4. package/adapters/claude/olko-github-pr/skills/lekker-review/references/pricing.json +9 -0
  5. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-consistency.md +76 -0
  6. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-conventions.md +158 -0
  7. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-implementation.md +76 -0
  8. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-quality.md +87 -0
  9. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-simplification.md +63 -0
  10. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-test-quality.md +172 -0
  11. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/phase4-comparison.md +42 -0
  12. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/profile.md +365 -0
  13. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/profiles/lekker-deep.md +83 -0
  14. package/adapters/claude/olko-github-pr/skills/lekker-review/references/revmux/profiles/lekker-medium.md +82 -0
  15. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/fixtures/context.json +3 -0
  16. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/fixtures/house-rules.md +11 -0
  17. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/fixtures/revmux-report.json +179 -0
  18. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/install-revmux-prompts.sh +66 -0
  19. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/revmux-adapter.mjs +261 -0
  20. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/revmux-engine.sh +156 -0
  21. package/adapters/claude/olko-github-pr/skills/lekker-review/scripts/selftest.mjs +76 -0
  22. package/package.json +1 -1
  23. package/plugins/olko-apple-kit/.claude-plugin/plugin.json +1 -1
  24. package/plugins/olko-creative/.claude-plugin/plugin.json +1 -1
  25. package/plugins/olko-garmin-kit/.claude-plugin/plugin.json +1 -1
  26. package/plugins/olko-git-tools/.claude-plugin/plugin.json +1 -1
  27. package/plugins/olko-github-pr/.claude-plugin/plugin.json +1 -1
  28. package/plugins/olko-github-pr/skills/lekker-review/SKILL.md +113 -4
  29. package/plugins/olko-github-pr/skills/lekker-review/fix-workflow.js +89 -7
  30. package/plugins/olko-github-pr/skills/lekker-review/references/agents/fix-verifier.md +56 -2
  31. package/plugins/olko-github-pr/skills/lekker-review/references/fix-mode.md +19 -1
  32. package/plugins/olko-github-pr/skills/lekker-review/references/pricing.json +9 -0
  33. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-consistency.md +76 -0
  34. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-conventions.md +158 -0
  35. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-implementation.md +76 -0
  36. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-quality.md +87 -0
  37. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-simplification.md +63 -0
  38. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/lenses/lekker-test-quality.md +172 -0
  39. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/phase4-comparison.md +42 -0
  40. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/profile.md +365 -0
  41. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/profiles/lekker-deep.md +83 -0
  42. package/plugins/olko-github-pr/skills/lekker-review/references/revmux/profiles/lekker-medium.md +82 -0
  43. package/plugins/olko-github-pr/skills/lekker-review/scripts/fixtures/context.json +3 -0
  44. package/plugins/olko-github-pr/skills/lekker-review/scripts/fixtures/house-rules.md +11 -0
  45. package/plugins/olko-github-pr/skills/lekker-review/scripts/fixtures/revmux-report.json +179 -0
  46. package/plugins/olko-github-pr/skills/lekker-review/scripts/install-revmux-prompts.sh +66 -0
  47. package/plugins/olko-github-pr/skills/lekker-review/scripts/revmux-adapter.mjs +261 -0
  48. package/plugins/olko-github-pr/skills/lekker-review/scripts/revmux-engine.sh +156 -0
  49. package/plugins/olko-github-pr/skills/lekker-review/scripts/selftest.mjs +76 -0
  50. package/plugins/olko-github-pr/skills/lekker-review/workflow.js +43 -6
  51. package/plugins/olko-obsidian/.claude-plugin/plugin.json +1 -1
  52. package/plugins/olko-product/.claude-plugin/plugin.json +1 -1
  53. package/plugins/olko-reflection/.claude-plugin/plugin.json +1 -1
  54. package/plugins/olko-release/.claude-plugin/plugin.json +1 -1
  55. package/plugins/olko-skill-meta/.claude-plugin/plugin.json +1 -1
  56. package/plugins/olko-web-ops/.claude-plugin/plugin.json +1 -1
@@ -0,0 +1,87 @@
1
+ ---
2
+ description: quality, security and data-integrity issues — Teifi's TS-1/TS-2 hard rules included
3
+ ---
4
+ ## Lens: lekker-quality
5
+
6
+ Review the change for quality, security, and data-integrity issues.
7
+
8
+ Axes to cover:
9
+ - Data Integrity: missing transactions on multi-step writes, optimistic
10
+ concurrency without locking, partial-failure with no rollback, silent data
11
+ loss in batch loops.
12
+ - Security: SQL/command injection, auth bypass, missing permission checks,
13
+ IDOR, secrets in logs, webhook signature not verified.
14
+ When the diff touches a controller, a route, or an auth decorator, do NOT hand-roll the
15
+ decorator grep. Use the `route-auth-map` skill, which already does exactly this and prints
16
+ method/path/controller/guards plus an explicit unguarded-endpoint section. If the Skill tool
17
+ is unavailable to you, read `~/.claude/skills/route-auth-map/SKILL.md` and follow its steps.
18
+ Compare the map before and after the diff: a route that gains a handler but no guard, or
19
+ loses `@Authenticated`/`@Permission`, is critical.
20
+ - Error Handling: unhandled promise rejections, empty catch blocks, missing
21
+ retries on transient failures, no dead-letter for failed jobs.
22
+ - Catch-block exit paths: when the diff adds or edits a branch inside a `catch`,
23
+ enumerate EVERY way control leaves that block - each early return, each
24
+ rethrow, and the fall-through - and say what the client sees on each. A new
25
+ branch that changes what gets logged or reported, while leaving a rethrow or
26
+ fall-through reachable for the same condition, is a finding: the error is now
27
+ silent AND still escapes. Do not accept "it reports correctly" as covering the
28
+ block; the report and the control flow are separate claims.
29
+ - Schema/Migration: NOT NULL without default, rename without two-step,
30
+ pgtyped queries invalidated, Prisma client out of sync.
31
+ - Naming/Typos: wrong casing convention, mixed conventions in same scope,
32
+ misspelled identifiers (these are bugs-in-waiting).
33
+ - Derived-value consistency: when the diff introduces a transform of some input
34
+ (normalise, trim, lowercase, parse, clamp, default), grep EVERY other use of
35
+ that raw input in the same scope and check they all go through the transform.
36
+ Half-applied transforms are a classic near-miss: the value is normalised at the
37
+ call site but the raw one is still used in a React dependency array, a cache
38
+ key, a log line, an equality check, or a second call site. Two spellings of the
39
+ "same" value then disagree. Enumerate the uses; do not eyeball the hunk.
40
+
41
+ ```bash
42
+ grep -n "<rawIdentifier>" <file> # every use, then confirm each is intended
43
+ ```
44
+
45
+ - Env vars: dead vars, renamed without migration, wrong fallback operator
46
+ (?? vs ||), type mismatch, leaked in logs.
47
+ - TypeScript type safety (TS-1): flag every type cast (`as X`, `<X>expr`)
48
+ and every use of `any`. Test files: only flag blatantly omitted types
49
+ (e.g. `any[]` on a clearly-typed list). All others: critical. Quote the
50
+ cast, explain the correct type, show the fix. Ask if they're Harry Potter.
51
+ Title the finding `[TS-1] ...`.
52
+ - No JavaScript files (TS-2): if the diff adds any `.js` file to a non-Liquid
53
+ theme repo, flag as critical — must be `.ts`. Title the finding `[TS-2] ...`.
54
+ - Dependency changes. Skip this axis entirely unless the diff touches
55
+ `package.json`, a lockfile, or a vendored dependency. Where it applies:
56
+ (a) A version bump is a behaviour change nobody in this PR wrote. If neither
57
+ the PR body nor a commit message cites the changelog or migration notes,
58
+ that is major: semver is a promise the maintainer may not have kept,
59
+ and a "patch" can carry a behavioural change.
60
+ (b) A bulk bump of several unrelated packages in one PR is major. When
61
+ it breaks the build you have lost which package did it. The fix is to
62
+ split it per package, or per genuinely related group.
63
+ (c) A `package.json` dependency change with no matching lockfile change in the
64
+ same diff, or a lockfile change with no `package.json` change and no
65
+ explanation, is critical: the lockfile is what actually ships.
66
+ (d) A new direct dependency that duplicates something already in the stack is
67
+ major. Name the existing thing that already solves it.
68
+ Raise NO naming, comment, complexity, or TS-1 finding inside a lockfile or a
69
+ `node_modules` path.
70
+
71
+ TS-1 and TS-2 are Teifi hard rules: their text is defined in full in `{{PROFILE}}`
72
+ (read it there before applying either). A finding for one of them MUST have its
73
+ title start with the bracketed tag, e.g. `[TS-1] ...` or `[TS-2] ...`, so the
74
+ caller can recognize it as a policy violation rather than an ordinary finding.
75
+
76
+ CI status, Sentry signals, and existing reviews may be present in `{{CONTEXT}}` —
77
+ read what is there before forming an opinion, and skip anything that is absent
78
+ rather than treating its absence as a finding.
79
+
80
+ Rules:
81
+ - Every finding must trace to a `+` line in the diff.
82
+ - Report file:line — description. No positive observations.
83
+ - Quote the verbatim offending line(s) from the diff — never paraphrased, never
84
+ reconstructed from memory — and give a concrete drop-in fix, or when the fix
85
+ is architectural, a minimal skeleton plus one sentence on what else must change.
86
+ - A finding you cannot quote and cannot fix is a finding you have not proven —
87
+ drop it instead of reporting it as a minor observation with no evidence.
@@ -0,0 +1,63 @@
1
+ ---
2
+ description: over-engineering and DRY violations, plus Teifi's debug-artifact hygiene table
3
+ ---
4
+ ## Lens: lekker-simplification
5
+
6
+ Review the change for over-engineering and DRY violations.
7
+
8
+ Look for:
9
+ - Copy-paste logic: identical blocks that differ only in a constant — flag
10
+ for extraction.
11
+ - Parallel implementations: two functions doing the same thing — one should
12
+ call the other.
13
+ - Unnecessary abstraction inversion: private helper called exactly once, adds
14
+ no reuse — should be inlined.
15
+ - Over-engineered control flow: nested ternaries / promise chains that could
16
+ be plain if/else or async/await.
17
+ - Config spread: same magic constant defined in multiple files.
18
+ - Debug artifacts and hygiene: apply the fixed severity table in §3 of
19
+ `{{PROFILE}}` (the Teifi conventions section) — `debugger` and
20
+ `.only`/`fit`/`fdescribe` are critical, an added `console.log`/`console.debug`
21
+ in production code and a hardcoded URL are major, an unreferenced
22
+ TODO/FIXME and a >3-line commented-out block are minor. Those severities
23
+ are fixed: do not soften them, and only flag occurrences the diff ADDED.
24
+ - Wrapper that adds nothing, factory for single implementation, layer-cake
25
+ anti-pattern (handler → service → repo with no logic in any layer).
26
+ - Feature flags always on/off, fallback that can never trigger, dual
27
+ implementations where old has no callers.
28
+ - Relocated complexity: a refactor that moves code without reducing the number
29
+ of concepts a reader must hold to follow it. Count them before and after. If
30
+ the count is unchanged, the restructuring did not simplify anything, and the
31
+ finding is that a cheaper move was available (deleting a branch, a mode, or a
32
+ layer outright, rather than re-centralising the same logic). Major when
33
+ the change is sold as a cleanup or refactor, minor otherwise.
34
+ - Feature logic in a shared module: feature-specific behaviour added to a
35
+ general-purpose util, a shared client, or a base class. The branch belongs in
36
+ the package that owns the concept. Name the owning layer in the fix.
37
+ - Dead code this diff orphans: when the diff replaces or reroutes something,
38
+ grep the worktree for remaining callers of what it superseded (the old helper,
39
+ the old component, a now-unreferenced constant, a flag that can no longer be
40
+ false). Enumerate what is now unreachable. NEVER propose a silent deletion:
41
+ the fix lists the orphans and asks the author to confirm removal.
42
+ Minor, or major when the dead path is still reachable from production code.
43
+
44
+ Only flag where duplication or complexity creates a real maintenance risk or
45
+ bug surface — not aesthetic preference.
46
+
47
+ When you flag a structural problem, name the move, not just the smell: replace a
48
+ chain of conditionals with a typed model or an explicit dispatcher, collapse
49
+ duplicate branches into one flow, separate orchestration from business logic,
50
+ move feature logic into the package that owns it, reuse the canonical helper
51
+ instead of a near-duplicate, delete a pass-through wrapper. Prefer the remedy
52
+ that removes moving pieces over one that spreads the same complexity around. A
53
+ finding that says "this is complex" without naming the restructuring is not
54
+ actionable: name the move or drop the finding.
55
+
56
+ Rules:
57
+ - Every finding must trace to a `+` line in the diff.
58
+ - Report file:line — description. No positive observations.
59
+ - Quote the verbatim offending line(s) — never paraphrased, never reconstructed
60
+ from memory — and give a concrete drop-in fix, or when the fix is
61
+ architectural, a minimal skeleton plus one sentence on what else must change.
62
+ - A finding you cannot quote and cannot fix is a finding you have not proven —
63
+ drop it instead.
@@ -0,0 +1,172 @@
1
+ ---
2
+ description: whether tests actually catch bugs — mutation-slip analysis, mock smell, Teifi test conventions
3
+ ---
4
+ ## Lens: lekker-test-quality
5
+
6
+ Review the change with a strict focus on test quality. This is NOT about
7
+ coverage numbers — it is about whether the tests actually catch bugs.
8
+
9
+ Step 1 — Inventory the tests. Read the diff at `{{SCOPE}}`. List every test
10
+ file/spec added or modified. If no test files are in the diff, note that and
11
+ continue to axis (f) below.
12
+
13
+ Step 2 — For each changed test file, read the full file from `{{WORKDIR}}`.
14
+
15
+ Step 3 — Evaluate each of these axes:
16
+
17
+ a) Meaningful assertions vs. smoke tests
18
+ - Does the test verify a specific outcome, or just that no exception was
19
+ thrown / a function returned truthy?
20
+ - Are assertions on the correct thing? (e.g. checking the return value vs
21
+ a side-effect that is the actual goal of the operation)
22
+ - Are there tautological assertions that are always true regardless of the
23
+ implementation? (e.g. `expect(true).toBe(true)`,
24
+ `expect(x).toBeDefined()` when x is always defined by construction)
25
+
26
+ b) Condition coverage
27
+ - Is the happy path tested?
28
+ - Are failure / error paths tested? (invalid input, external API error,
29
+ empty collection, null/undefined, 0 or negative numbers, pagination edge
30
+ cases, missing env vars)
31
+ - For every new if/else or switch in the changed business logic, is each
32
+ branch exercised by at least one test case?
33
+
34
+ c) Regression tests
35
+ - If the change fixes a bug (the ticket mentions a bug, or the diff
36
+ contains "fix" language), is there a regression test that would have
37
+ caught the original bug? Flag if absent.
38
+ - If the change adds a new feature, are the known edge cases of that feature
39
+ tested?
40
+
41
+ d) Mutation-slip analysis (mental mutation testing)
42
+ For the most critical assertions in the test suite, ask: would a simple
43
+ mutation in the production code slip through undetected?
44
+
45
+ Consider these mutation classes:
46
+ - Off-by-one: `> N` changed to `>= N`
47
+ - Operator flip: `&&` to `||`, `===` to `!==`
48
+ - Missing null/undefined guard: remove a `?? default`
49
+ - Wrong variable: using `a` where `b` was intended
50
+ - Return-value swap: returning the wrong field from an object
51
+ - Early-return removed: a guard clause deleted
52
+
53
+ For each mutation class relevant to the changed business logic, determine
54
+ whether at least one test assertion would catch it. Summarize as a short
55
+ paragraph: "Mutations that would slip through: ..." or "No obvious
56
+ mutation-slip gaps found."
57
+
58
+ e) Test isolation and reliability
59
+ - Do tests share mutable state across cases without resetting between runs
60
+ (a beforeEach that does not clean up)?
61
+ - Are there tests that depend on execution order or global singletons?
62
+ - Could a test make a real network/DB call in CI (flaky)? The fix is a
63
+ simple injected fake at the boundary, not blanket module mocking — see
64
+ axis (g).
65
+ - Are async tests properly awaited? (floating promises, missing `await` on
66
+ `expect().resolves`, unhandled rejections)
67
+
68
+ f) Test-to-code ratio signal
69
+ If the diff adds > 50 lines of new business logic with zero new or modified
70
+ test files, flag it explicitly. Then check whether existing test files
71
+ already cover the new code paths:
72
+ ```bash
73
+ find <workdir> -name "*.test.ts" -o -name "*.spec.ts" | \
74
+ xargs grep -l "<key changed symbol>" 2>/dev/null | head -5
75
+ ```
76
+ Report whether existing coverage closes the gap or not.
77
+
78
+ g) Mock smell — test the behavior, not the way it's built
79
+ (House standard: Notion "To mock or not to mock" —
80
+ https://app.notion.com/p/teifi/To-mock-or-not-to-mock-36ff8ed0f7db80f09495d174e3b86cd6)
81
+ Flag tests coupled to *how* the code is built rather than *what* it does for
82
+ the user. For each smell, do NOT just criticize: give the concrete no-mock
83
+ refactor. The default fix is almost always the same shape — pull the logic
84
+ into a pure function (functional core) and test that directly, leaving a thin
85
+ shell covered by a few real integration tests.
86
+
87
+ - Mocking IO just to reach logic: a GraphQL/fetch/DB client mocked only so a
88
+ test can read a computed value back. It couples to the query shape AND the
89
+ markup while barely testing the logic, and silently rots as the real
90
+ dependency drifts from the mock. This applies just as much on the backend
91
+ as in a component: mocking a service client (e.g. `vi.mocked(epicorClient)`
92
+ returning a canned async generator/array) just to reach a reassembly loop
93
+ or a multi-stream join is the same smell as mocking `fetch` in a component.
94
+ → Extract the computation into a pure function over plain data and test
95
+ that with literals (no mocks, no render, no mocked client). Cover
96
+ fetch-and-wire once with a real integration test. For streaming/paging
97
+ code specifically: a shared `collect()`/reassembly helper should be a
98
+ pure function of `AsyncGenerator<T[]> → Promise<T[]>` (or similar) fed a
99
+ plain fake generator in its own test — never a mocked client — and any
100
+ call-site logic that combines multiple streams (parallel joins, chunked
101
+ batching) should be its own pure function tested the same way.
102
+ - Spying on calls / asserting call shape: `toHaveBeenCalledWith`,
103
+ `toHaveBeenCalledTimes`, a `vi.fn()` used as a probe. Asserts *how* a
104
+ function was called, freezing batching/page-size/call-count in place even
105
+ when the output is identical.
106
+ → Assert the output, never the call log. If a boundary is genuinely
107
+ needed, inject a simple fake (a plain function returning canned data),
108
+ not a spy.
109
+ - Mocking a component to read its props back (re-emitting props as `data-*`
110
+ attributes, then asserting on them): the assertions are about the mock, and
111
+ break on a component swap or prop rename that changes nothing a user sees.
112
+ → Pull the logic (pagination, display state) into a pure function over
113
+ plain values — e.g. `paginate(items, page, pageSize)` — and assert the
114
+ value it returns.
115
+ - Mocking a query hook (`useQuery` / a `use-X` hook) to hand a component
116
+ canned data: re-tests React Query's own plumbing and couples to the hook's
117
+ return shape.
118
+ → Keep the `queryFn` as IO (integration-tested); move the transform into a
119
+ pure function wired through React Query's `select` option, and unit-test
120
+ that pure function with plain data.
121
+ - Asserting implementation details: that a component is memoized, that a
122
+ specific child renders, that work happens through a fixed sequence of
123
+ calls.
124
+ → Delete the assertion; assert the user-visible behavior instead.
125
+
126
+ When a mock IS the right call, do NOT flag it: a real IO boundary that must
127
+ be exercised where a fake is impractical, non-determinism that must be pinned
128
+ (time, randomness, injected network failure), or a dependency that genuinely
129
+ cannot run in the test environment. The tell for a good mock: the behavior
130
+ under test only exists because the boundary did something (e.g. a retry
131
+ banner that appears only when the fetch rejects). Even then, prefer an
132
+ injected simple fake over a module-level spy, and assert what the user sees.
133
+
134
+ h) Teifi test conventions
135
+ Read `{{PROFILE}}` and apply its §4 (test conventions) to every test file in
136
+ the diff: `__tests__/` placement beside the source, `.test.ts` vs
137
+ `.test.tsx`, English-sentence test names, `describe` nesting <= 2,
138
+ `afterEach(cleanup)`, `retry: false` on the test QueryClient (major — its
139
+ absence hangs the suite on a failing query), the getByRole → getByLabelText
140
+ → getByText → getByTestId priority, module-boundary mocking, `safeParse` for
141
+ Zod schemas, and the typed `it.each` matrix with AC tags plus the
142
+ exhaustiveness assertion whenever behaviour depends on 3+ independent inputs.
143
+ A `vi.mock('@lib/common/…')` inside `common/` is major, not style: the
144
+ alias resolves only from `web/`, so the mock silently does nothing and the
145
+ test passes against the real module. Quote the line, give the relative path
146
+ as the fix.
147
+ A comment explaining WHY a test exists (a captured past bug, a subtle
148
+ contract) is sanctioned — never flag it as over-commenting.
149
+
150
+ Report each per-line issue as:
151
+ `test-file:line — <concise description of the gap or weakness>`
152
+
153
+ For every mock-smell finding from axis (g), append the no-mock fix on the next
154
+ line as `→ Fix: <pure-function / simple-fake refactor in one line>`. A
155
+ criticism without a fix is incomplete.
156
+
157
+ Report the mutation-slip analysis as a single paragraph under a
158
+ "**Mutation-slip risk:**" heading — not as line items.
159
+
160
+ Rules:
161
+ - Every per-line finding must trace to test code in the diff, OR to business
162
+ logic added in the diff that has no test coverage at all.
163
+ - Report problems only. No praise for tests that meet the bar.
164
+ - If no test files are changed AND no existing tests cover the new code paths,
165
+ report: "No test coverage for new code paths."
166
+ - Quote the verbatim offending line(s) — never paraphrased, never reconstructed
167
+ from memory. When the finding IS the absence of a test, there is no test line
168
+ to quote: quote the untested production line(s) from the diff instead, and
169
+ give the test that should exist as the fix. Never drop a "no coverage"
170
+ finding just because nothing bad is written down — absence is the finding.
171
+ - A finding you cannot quote and cannot fix (outside the no-coverage case
172
+ above) is a finding you have not proven — drop it instead.
@@ -0,0 +1,42 @@
1
+ # Phase 4: revmux engine on real PRs
2
+
3
+ One row per live `--engine revmux` run. Numbers come from revmux `stats` via
4
+ `scripts/revmux-adapter.mjs`; USD is the `pricing.json` list-price placeholder
5
+ (`verified: false`, output price applied to every token, synth+verify unpriced),
6
+ and every run so far went through the Max subscription, so USD is a relative
7
+ signal only. "Confirmed" means the main loop kept the finding after reading the
8
+ worktree; "merged" means two revmux findings described one mechanism.
9
+
10
+ | Date | PR | Depth | Wall | Tokens | USD est. | revmux findings | Kept | Re-severity | Hard-rule re-promotions | Merged | False positives |
11
+ |---|---|---|---|---|---|---|---|---|---|---|---|
12
+ | 2026-09-09 | evi-integrations #544 | medium (auto) | 15m40s | 12,604,506 | $169.53 | 4 | 3 (1 C / 2 I) + obs | 1 important → critical (lost role grants, watermark written before post-pass) | 0 | 0 | 0 |
13
+ | 2026-09-09 | evi-integrations #537 | deep (auto, *.sql) | 11m08s | 11,245,064 | $141.61 | 4 | 3 (0 C / 2 I / 1 idiomatic) + obs | none | 0 | 1 (#1 reactivation + #4 guard removal → one Important) | 0 |
14
+
15
+ ## Notes per run
16
+
17
+ ### #544
18
+
19
+ - revmux under-rated the watermark ordering bug as important; the main loop
20
+ raised it to Critical after tracing `executeSyncStrategy` writing the CUSTOMER
21
+ watermark before the role post-pass.
22
+ - One finding referenced `getShopFlag`, which does not exist in this repo;
23
+ rewritten to an env toggle with Reflag as the policy target. Class: fix code
24
+ invented from another repo's helper. Worth a lens note.
25
+ - Depth medium, so verify ran on Criticals only; revmux's own verify stage
26
+ covered all four.
27
+
28
+ ### #537
29
+
30
+ - Two of four findings were the same mechanism seen from two lenses (bugs,
31
+ adversarial). Synthesis did not merge them; the main loop did.
32
+ - All four confirmed against the worktree. The idiomatic one (comment
33
+ narrating deleted code) is exactly the teifi-conventions §2 class.
34
+ - Reactivation and soft-delete files are outside the diff, so the main
35
+ finding could only be anchored on the sync-path signal it replaces.
36
+
37
+ ## Still open
38
+
39
+ - Same PRs through the default `workflow` engine for a side-by-side (Oleg-driven).
40
+ - Read-only enforcement: revmux default `--tools` includes Bash; prompt-enforced
41
+ only. `--tools=Read,Grep,Glob,WebFetch,WebSearch` override not yet applied.
42
+ - Adapter: price synth + verify once revmux reports a per-model split.