@iceinvein/agent-skills 0.1.37 → 0.1.39

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. package/package.json +1 -1
  2. package/skills/bounded-context-auditor/SKILL.md +15 -3
  3. package/skills/bounded-context-auditor/skill.json +1 -1
  4. package/skills/codebase-architecture/SKILL.md +4 -4
  5. package/skills/codebase-architecture/skill.json +1 -1
  6. package/skills/cognitive-load-auditor/SKILL.md +11 -9
  7. package/skills/cognitive-load-auditor/skill.json +1 -1
  8. package/skills/cohesion-analyzer/SKILL.md +1 -1
  9. package/skills/cohesion-analyzer/skill.json +1 -1
  10. package/skills/composability-auditor/SKILL.md +5 -5
  11. package/skills/composability-auditor/skill.json +1 -1
  12. package/skills/contract-enforcer/SKILL.md +5 -5
  13. package/skills/contract-enforcer/skill.json +1 -1
  14. package/skills/coupling-auditor/SKILL.md +2 -2
  15. package/skills/coupling-auditor/skill.json +1 -1
  16. package/skills/cover-letter/SKILL.md +18 -20
  17. package/skills/cover-letter/skill.json +7 -2
  18. package/skills/cover-letter-audit/SKILL.md +20 -20
  19. package/skills/cover-letter-audit/skill.json +7 -2
  20. package/skills/cover-letter-persona/SKILL.md +13 -13
  21. package/skills/cover-letter-persona/skill.json +7 -2
  22. package/skills/cover-letter-rewrite/SKILL.md +18 -16
  23. package/skills/cover-letter-rewrite/skill.json +7 -2
  24. package/skills/cover-letter-write/SKILL.md +25 -20
  25. package/skills/cover-letter-write/skill.json +7 -2
  26. package/skills/cqs-auditor/SKILL.md +19 -47
  27. package/skills/cqs-auditor/skill.json +1 -1
  28. package/skills/demeter-enforcer/SKILL.md +5 -5
  29. package/skills/demeter-enforcer/skill.json +1 -1
  30. package/skills/dependency-direction-auditor/SKILL.md +1 -1
  31. package/skills/dependency-direction-auditor/skill.json +1 -1
  32. package/skills/design-review/SKILL.md +6 -2
  33. package/skills/design-review/skill.json +1 -1
  34. package/skills/error-strategist/SKILL.md +3 -3
  35. package/skills/error-strategist/skill.json +1 -1
  36. package/skills/event-design-reviewer/SKILL.md +3 -3
  37. package/skills/event-design-reviewer/skill.json +1 -1
  38. package/skills/evolution-analyzer/SKILL.md +4 -3
  39. package/skills/evolution-analyzer/skill.json +1 -1
  40. package/skills/gestalt-reviewer/SKILL.md +8 -4
  41. package/skills/gestalt-reviewer/skill.json +1 -1
  42. package/skills/idempotency-guardian/SKILL.md +6 -6
  43. package/skills/idempotency-guardian/skill.json +1 -1
  44. package/skills/improve-my-codebase/CATALOGUE-FIELDS.md +2 -2
  45. package/skills/improve-my-codebase/SKILL.md +68 -27
  46. package/skills/improve-my-codebase/skill.json +1 -1
  47. package/skills/index.json +33 -33
  48. package/skills/integration-pattern-auditor/SKILL.md +2 -2
  49. package/skills/integration-pattern-auditor/skill.json +1 -1
  50. package/skills/magpie/README.md +3 -5
  51. package/skills/magpie/SKILL.md +39 -536
  52. package/skills/magpie/package.json +1 -1
  53. package/skills/magpie/references/critic.md +58 -0
  54. package/skills/magpie/references/peer-review.md +84 -0
  55. package/skills/magpie/references/specialists.md +391 -0
  56. package/skills/magpie/scripts/__tests__/helper.test.ts +40 -0
  57. package/skills/magpie/scripts/__tests__/skill-lint.test.ts +116 -28
  58. package/skills/magpie/scripts/__tests__/status-cmd.test.ts +13 -0
  59. package/skills/magpie/scripts/helper.js +24 -13
  60. package/skills/magpie/scripts/status-cmd.ts +10 -1
  61. package/skills/magpie/skill.json +2 -1
  62. package/skills/module-secret-auditor/SKILL.md +8 -5
  63. package/skills/module-secret-auditor/skill.json +1 -1
  64. package/skills/port-adapter-auditor/SKILL.md +3 -3
  65. package/skills/port-adapter-auditor/skill.json +1 -1
  66. package/skills/rams-design-audit/SKILL.md +4 -2
  67. package/skills/rams-design-audit/skill.json +1 -1
  68. package/skills/seam-finder/SKILL.md +2 -2
  69. package/skills/seam-finder/skill.json +1 -1
  70. package/skills/simplicity-razor/SKILL.md +4 -4
  71. package/skills/simplicity-razor/skill.json +1 -1
  72. package/skills/temporal-coupling-detector/SKILL.md +2 -2
  73. package/skills/temporal-coupling-detector/skill.json +1 -1
  74. package/skills/terse/SKILL.md +12 -7
  75. package/skills/terse/skill.json +1 -1
  76. package/skills/type-driven-designer/SKILL.md +8 -8
  77. package/skills/type-driven-designer/skill.json +1 -1
  78. package/skills/unidirectional-flow-enforcer/SKILL.md +2 -2
  79. package/skills/unidirectional-flow-enforcer/skill.json +1 -1
@@ -18,7 +18,7 @@ A visual perception audit based on the Gestalt principles of visual organization
18
18
  - After generating multi-element layouts (forms, dashboards, cards, navigation)
19
19
  - When spacing between elements feels arbitrary or inconsistent
20
20
 
21
- **Not for:** Color theory, typography selection, interaction design, animation, or accessibility compliance. This skill evaluates *spatial organization and visual grouping*, not visual styling.
21
+ **Not for:** Color theory, typography selection, interaction design, animation, or accessibility compliance. This skill evaluates *spatial organization and visual grouping*, not visual styling. For decoration and element-earning-its-place questions use `rams-design-audit`; for option overload and memory demands use `cognitive-load-auditor` — the three compose on a full UI review.
22
22
 
23
23
  ## The Process
24
24
 
@@ -62,9 +62,9 @@ If any two adjacent spacings are the same, the grouping is ambiguous.
62
62
 
63
63
  **The consistency test:** If two elements look the same, they should behave the same. If they behave differently, they must look different.
64
64
 
65
- ### 3. Closure — Can the mind complete the shapes?
65
+ ### 3. Closure & Common Region — Can the mind complete the shapes, and do enclosures group?
66
66
 
67
- **Principle:** The visual system completes incomplete shapes and perceives enclosed regions as groups. You don't always need explicit borders — implied boundaries work.
67
+ **Principle:** Two related principles audited together. Closure (classic Gestalt): the visual system completes incomplete shapes, so partial elements must read as deliberately incomplete, not broken. Common region (Palmer's extension): elements inside the same enclosed region are perceived as a group, so you don't always need explicit borders — implied boundaries work.
68
68
 
69
69
  **Audit checklist:**
70
70
  - Are containers and regions perceivable without heavy borders?
@@ -96,7 +96,7 @@ If any two adjacent spacings are the same, the grouping is ambiguous.
96
96
  - Vertical rhythm breaks — consistent spacing except one section that's arbitrarily different
97
97
  - Centered content inside left-aligned containers — creates an unstable axis
98
98
 
99
- **The squint test:** Squint at the layout (or blur it). The alignment axes should still be visible. If the structure disappears when you can't read the text, the visual organization depends on reading, not perception.
99
+ **The squint test:** Squint at the layout (or blur it). The alignment axes should still be visible. If the structure disappears when you can't read the text, the visual organization depends on reading, not perception. (Agent equivalent: with a screenshot, downscale/blur it and check that groupings still read; source-only, verify shared alignment values instead — common margins, grid columns, consistent spacing tokens.)
100
100
 
101
101
  ### 5. Figure-Ground — Is foreground clear?
102
102
 
@@ -144,6 +144,8 @@ GESTALT: Contact form
144
144
 
145
145
  Decision engine. After generating UI with multiple elements, the agent audits the spatial organization against Gestalt principles. The audit identifies specific violations and provides concrete spacing/alignment/styling fixes. The human sees the audit alongside the layout.
146
146
 
147
+ **Acquiring the UI when auditing something the agent didn't just write:** this is the most rendering-dependent of the UI audits. Prefer a rendered view (user screenshot, or a browser/screenshot tool if available). Source-only, audit the measurable proxies (spacing values, alignment tokens, border/shadow styles) and say the audit ran source-only.
148
+
147
149
  The agent prioritizes **proximity** (most impactful, most commonly violated) and **continuity** (alignment issues are the most visually jarring). Similarity and figure-ground are audited when relevant but are less frequently the root cause of "this layout feels off."
148
150
 
149
151
  ## Guard Rails
@@ -167,3 +169,5 @@ The agent prioritizes **proximity** (most impactful, most commonly violated) and
167
169
  | **Figure-Ground** | Foreground separates from background | Is primary content unambiguously the figure? |
168
170
  | **Common Fate** | Elements moving together are grouped | Do animations group the right elements? |
169
171
  | **Prägnanz** | Mind prefers the simplest interpretation | Is the simplest reading of the layout the correct one? |
172
+
173
+ (Common Fate and Prägnanz are reference-only: audit Common Fate when the UI has animation or motion; Prägnanz is the summary lens behind the other principles, not a separate report line.)
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "gestalt-reviewer",
3
- "version": "1.0.1",
3
+ "version": "1.0.2",
4
4
  "description": "Gestalt-inspired visual perception audit: proximity, similarity, closure, continuity, and figure-ground analysis for UI layouts",
5
5
  "author": "iceinvein",
6
6
  "type": "prompt",
@@ -1,11 +1,11 @@
1
1
  ---
2
2
  name: idempotency-guardian
3
- description: Use when designing or reviewing API endpoints that mutate state, reviewing event/message handlers, when retry logic exists anywhere in the system, when debugging duplicate side effects (double charges, duplicate emails, duplicated records), or when designing webhook/callback handlers. Trigger on "why did the customer get charged twice?", "the queue redelivered and now we have duplicate rows", or when writing handlers for external input. NOT for read-only operations, pure functions, or coupling/dependency structure.
3
+ description: Use when designing or reviewing API endpoints that mutate state, event/message handlers, or webhook handlers, when retry logic exists anywhere in the system, or when debugging duplicate side effects (double charges, duplicate emails). Trigger on "why did the customer get charged twice?" or "the queue redelivered and now we have duplicate rows". NOT for read-only operations, pure functions, or coupling/dependency structure.
4
4
  ---
5
5
 
6
6
  # Idempotency Guardian
7
7
 
8
- A distributed systems analysis framework based on Pat Helland's *Idempotence Is Not a Medical Condition* (2012), the HTTP specification (RFC 7231), and messaging literature (Hohpe & Woolf's *Enterprise Integration Patterns*). Networks are unreliable. Messages get retried, consumers process the same event twice, payments get resubmitted, webhooks fire multiple times. The fundamental question is never "will this be called twice?" — it will be. The real question is: "when this is called twice with the same input, will the system still be correct?" An idempotent operation produces the same result whether executed once or multiple times. Idempotency is not an optimization or a nice-to-have — it's the minimum safety property for anything that can be called more than once. Without it, retries cause data corruption, duplicate side effects, and cascading failures.
8
+ A distributed systems analysis framework based on Pat Helland's *Idempotence Is Not a Medical Condition* (2012), the HTTP specification (RFC 9110), and messaging literature (Hohpe & Woolf's *Enterprise Integration Patterns*). Networks are unreliable. Messages get retried, consumers process the same event twice, payments get resubmitted, webhooks fire multiple times. The fundamental question is never "will this be called twice?" — it will be. The real question is: "when this is called twice with the same input, will the system still be correct?" An idempotent operation produces the same result whether executed once or multiple times. Idempotency is not an optimization or a nice-to-have — it's the minimum safety property for anything that can be called more than once. Without it, retries cause data corruption, duplicate side effects, and cascading failures.
9
9
 
10
10
  **Core principle:** An idempotent operation produces the same observable result whether executed once or executed multiple times with the same input. Idempotency is the only reliable way to handle retries, at-least-once delivery, and network unreliability. State changes should be idempotent; side effects require explicit protection.
11
11
 
@@ -177,7 +177,7 @@ Mutation: POST /api/orders
177
177
  1. Insert order in database → Protected (inside transaction, only runs if idempotency key check passes)
178
178
  2. Emit OrderCreated event → UNPROTECTED (if request is retried, event fires twice; downstream processes duplicate order)
179
179
  3. Send confirmation email → UNPROTECTED (if request is retried, email sent twice)
180
- Fix: Emit event and send email inside the transaction, or defer with exactly-once guarantee
180
+ Fix: Transactional outbox — write outbox records for the event and the email in the same transaction as the order insert; a relay delivers them with deduplication, so a retried request never fires the side effects twice
181
181
 
182
182
  Mutation: ProcessPaymentHandler (event consumer)
183
183
  Side effects:
@@ -216,7 +216,7 @@ MUTATION: POST /api/orders (UNHEALTHY)
216
216
  Side effects: 1. Insert order in DB, 2. Emit OrderCreated event, 3. Send confirmation email
217
217
  Side effect safety: 1. DB (duplicate row), 2. Event (fires twice), 3. Email (sent twice)
218
218
  Risk: Retry sends confirmation email twice, publishes event twice (downstream processes order twice), inserts duplicate row
219
- Fix: Add Idempotency-Key header check (atomic check-and-cache). Move event and email into same transaction as order insert. Or: use event sourcing + dedup consumer.
219
+ Fix: Add Idempotency-Key header check (atomic check-and-cache). Write the event and the email as outbox records in the same transaction as the order insert (transactional outbox); a relay delivers them to idempotent consumers. Or: use event sourcing + dedup consumer.
220
220
  ```
221
221
 
222
222
  ```
@@ -246,13 +246,13 @@ Decision engine, prioritizes by blast radius (financial > external API > data in
246
246
 
247
247
  ## Guard Rails
248
248
 
249
- **State changes should be idempotent. Side effects require explicit protection.** Don't confuse them. Making an INSERT idempotent via upsert is straightforward. Making sure email is only sent once requires additional work (dedup, guard checks, or moving email into the same transaction).
249
+ **State changes should be idempotent. Side effects require explicit protection.** Don't confuse them. Making an INSERT idempotent via upsert is straightforward. Making sure email is only sent once requires additional work (dedup, guard checks, or the outbox pattern: record the send intent transactionally, deliver via a relay).
250
250
 
251
251
  **"Just deduplicate" is not sufficient.** Deduplication prevents re-execution of the operation but not of its side effects. If you've already called Stripe.charge() and saved the result, dedup prevents you from calling it again — but only if you deduplicate before calling the external API.
252
252
 
253
253
  **Idempotency keys need lifecycle management.** You can't keep them forever. Set an expiry (24-72 hours is typical). After expiry, the same key can be reused. This is fine because the operation likely won't be retried after that time window.
254
254
 
255
- **Don't confuse idempotent with safe.** An operation can be idempotent but still unsafe (e.g., if it reads stale data). Idempotency is about repeated execution of the operation, not about correctness under concurrent execution.
255
+ **Don't confuse idempotent with correct-under-concurrency.** An operation can be idempotent but still race (e.g., if it reads stale data between check and write). Idempotency is about repeated execution of the operation, not about correctness under concurrent execution. (And note "safe" in the HTTP spec is a separate term of art meaning side-effect-free.)
256
256
 
257
257
  **Distributed side effects are the hardest part.** If your operation calls multiple external systems (Stripe + SendGrid + analytics), protecting all of them is complex. Options: use outbox pattern (record intent, external daemon delivers), use orchestration (record state machine, step through atomically), or accept some risk.
258
258
 
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "idempotency-guardian",
3
- "version": "1.0.1",
3
+ "version": "1.0.2",
4
4
  "description": "Helland-inspired idempotency analysis: classify mutation points, check protection mechanisms, evaluate side effect safety for retry-safe systems",
5
5
  "author": "iceinvein",
6
6
  "type": "prompt",
@@ -1,6 +1,6 @@
1
1
  # Catalogue fields for improve-my-codebase
2
2
 
3
- The `improve-my-codebase` orchestrator reads `skills/index.json` and routes audits based on two fields:
3
+ The `improve-my-codebase` orchestrator routes audits from the inline routing table in its SKILL.md (so routing works after installation, where `skills/index.json` is not present). The table mirrors two fields maintained in `skills/index.json`; when changing `applies`/`quick` for any audit, update both places:
4
4
 
5
5
  ## `applies: string[]`
6
6
 
@@ -12,7 +12,7 @@ Areas a skill is relevant to. The orchestrator runs an audit only when at least
12
12
  | `ui` | Requires UI surface | `.tsx`/`.jsx`/`.vue`/`.svelte`/`.html`/`.css` files, or `react`/`vue`/`svelte`/`solid` in `package.json` deps |
13
13
  | `domain` | Requires domain layer | Directory named `domain/`, `entities/`, `aggregates/`, or domain-driven naming |
14
14
  | `integration` | Requires messaging/events | Directory named `events/`, `messaging/`, `queues/`, or deps `kafkajs`/`amqplib`/`bullmq` |
15
- | `architecture` | Cross-cutting structural | Always matches in projects with > 5 files |
15
+ | `architecture` | Cross-cutting structural | Always matches in projects with >= 5 files |
16
16
  | `errors` | Error handling concerns | Always matches in any non-trivial codebase |
17
17
  | `legacy` | Modifying existing code | Only fires when invoked with `diff` mode |
18
18
 
@@ -25,7 +25,7 @@ A meta-orchestrator that runs the existing audit skills in this package across a
25
25
  ```
26
26
  /improve-my-codebase # full sweep
27
27
  /improve-my-codebase quick # fast subset, top 5 issues
28
- /improve-my-codebase diff # only changed files vs. main
28
+ /improve-my-codebase diff # only changed files vs. the base branch
29
29
  /improve-my-codebase interactive # interview-driven
30
30
  /improve-my-codebase focus <area> # narrow to one axis
31
31
  /improve-my-codebase module <path> # scope to a directory or file
@@ -36,8 +36,7 @@ A meta-orchestrator that runs the existing audit skills in this package across a
36
36
  Argument rules:
37
37
  - Positional. No leading `--`.
38
38
  - Modes: `quick`, `diff`, `interactive`. First match wins; mutually exclusive.
39
- - Scope filters: `focus <area>`, `module <path>`. Compose freely with each other and with default/diff modes.
40
- - `quick` and `interactive` ignore scope tokens (warn, do not error).
39
+ - Scope filters: `focus <area>`, `module <path>`. Compose freely with each other and with every mode, including `quick` and `interactive` (interactive uses provided scope values as the defaults for its questions).
41
40
  - Unknown tokens: warn and drop, do not abort.
42
41
 
43
42
  ## Process Overview
@@ -77,13 +76,14 @@ The orchestrator receives a single string of positional args after the command.
77
76
 
78
77
  **Rules:**
79
78
  - Empty args produce `{mode: "full", scope: {module: null, focus: null}}`.
80
- - `quick` or `interactive` with scope tokens present: keep mode, set scope to `{module: null, focus: null}`, emit warning "scope tokens ignored in <mode> mode".
79
+ - Scope tokens are honored in every mode. `quick focus ui` runs the quick subset narrowed to `ui`.
81
80
  - `focus` without a value: emit error listing valid areas (from catalogue), exit before dispatch.
82
81
  - `module` without a value: emit error "module requires a path argument", exit before dispatch.
83
82
  - `module <path>` where path does not exist: emit error "no such module: <path>", exit before dispatch.
84
83
 
85
84
  **Valid focus areas** (derived from the `applies` vocabulary):
86
- `any`, `ui`, `domain`, `integration`, `architecture`, `errors`, `legacy`.
85
+ `ui`, `domain`, `integration`, `architecture`, `errors`, `legacy`.
86
+ (`any` is not a focus area: audits with `applies: any` always run, so there is nothing to narrow to. `focus any` gets the standard invalid-area error.)
87
87
 
88
88
  **Worked examples:**
89
89
 
@@ -95,7 +95,7 @@ The orchestrator receives a single string of positional args after the command.
95
95
  | `"focus architecture"` | `{mode: "full", scope: {module: null, focus: "architecture"}}` |
96
96
  | `"diff focus architecture"` | `{mode: "diff", scope: {module: null, focus: "architecture"}}` |
97
97
  | `"module src/auth focus architecture"` | `{mode: "full", scope: {module: "src/auth", focus: "architecture"}}` |
98
- | `"quick focus ui"` | `{mode: "quick", scope: {module: null, focus: null}}` plus warning |
98
+ | `"quick focus ui"` | `{mode: "quick", scope: {module: null, focus: "ui"}}` |
99
99
  | `"banana"` | `{mode: "full", scope: {module: null, focus: null}}` plus warning |
100
100
 
101
101
  ## Phase 2: Detect stack
@@ -149,11 +149,48 @@ Choose which audits to dispatch.
149
149
  **Inputs:**
150
150
  - Detection result from Phase 2.
151
151
  - Parsed args from Phase 1.
152
- - Catalogue: `skills/index.json`.
152
+ - Catalogue: the inline routing table below.
153
+
154
+ **The catalogue (routing table):**
155
+
156
+ This table is the catalogue; no external file is needed to route. (In the source repo it mirrors the `applies`/`quick` fields in `skills/index.json`; the repo's audit script keeps them in sync.)
157
+
158
+ | Audit | Applies | Quick |
159
+ |-------|---------|-------|
160
+ | bounded-context-auditor | domain | no |
161
+ | codebase-architecture | any, architecture | no |
162
+ | cognitive-load-auditor | ui | no |
163
+ | cohesion-analyzer | architecture | yes |
164
+ | complexity-accountant | any | no |
165
+ | composability-auditor | any | no |
166
+ | contract-enforcer | any | no |
167
+ | coupling-auditor | architecture | yes |
168
+ | cqs-auditor | architecture | yes |
169
+ | demeter-enforcer | architecture | yes |
170
+ | dependency-direction-auditor | architecture | yes |
171
+ | design-review | any | no |
172
+ | error-strategist | errors | no |
173
+ | event-design-reviewer | integration, domain | no |
174
+ | evolution-analyzer | any | no |
175
+ | gestalt-reviewer | ui | yes |
176
+ | idempotency-guardian | integration | no |
177
+ | integration-pattern-auditor | integration | no |
178
+ | module-secret-auditor | architecture | yes |
179
+ | port-adapter-auditor | architecture | no |
180
+ | rams-design-audit | ui | yes |
181
+ | seam-finder | legacy | yes |
182
+ | simplicity-razor | any | no |
183
+ | temporal-coupling-detector | any | yes |
184
+ | type-driven-designer | any | no |
185
+ | unidirectional-flow-enforcer | ui | no |
186
+
187
+ 26 audits total (`catalogue_audit_count = 26`).
188
+
189
+ **Locating audit content:** each routed audit's principles are read from its sibling skill directory, resolved relative to THIS file's location: the parent directory of `improve-my-codebase/` is the skills root (`skills/` in the source repo, `.claude/skills/` when installed). An audit's content is at `<skills-root>/<audit-id>/SKILL.md`. If that file does not exist (the audit is not installed), skip the audit, record `{audit, status: "not-installed"}`, and continue; suggest installing it in the terminal summary.
153
190
 
154
191
  **Algorithm:**
155
192
 
156
- 1. Load catalogue from `skills/index.json`. If the file is missing, unreadable, or fails JSON parse, hard fail per the error table (the orchestrator cannot route without a catalogue). Filter to entries with both `applies` and `quick` fields (i.e. audit skills).
193
+ 1. Start from the routing table above.
157
194
  2. **Apply detection filter:** keep an audit if any of its `applies` values matches an active signal:
158
195
  - `any` always matches.
159
196
  - `ui` matches if `hasUI`.
@@ -174,6 +211,8 @@ Choose which audits to dispatch.
174
211
 
175
212
  **Module scope** is **not** applied here. It carries forward to Phase 4 as a per-subagent scope hint, because audit routing is the same regardless of whether the audit examines the whole repo or one path.
176
213
 
214
+ 6. For each surviving audit, verify its SKILL.md exists at `<skills-root>/<audit-id>/SKILL.md` (see Locating audit content above). Missing ones are recorded as `not-installed` and removed from the dispatch list; if that empties the list, abort with "Matched audits are not installed: <list>. Install them and rerun."
215
+
177
216
  **Worked example:**
178
217
  - Repo: Bun backend with `events/` directory.
179
218
  - Detection: `{hasUI: false, hasDomainLayer: false, hasIntegration: true, hasArchitecture: true}`.
@@ -186,10 +225,11 @@ Choose which audits to dispatch.
186
225
 
187
226
  When `mode === "diff"`, before dispatch the orchestrator computes the changed-file set:
188
227
 
189
- 1. Run `git diff --name-only origin/main...HEAD` (fall back to `git diff --name-only main...HEAD` if no `origin` remote).
190
- 2. Filter out paths that no longer exist (deleted files) and paths in the standard ignored set (`node_modules/`, `.git/`, `dist/`, `build/`, `coverage/`, `tests/fixtures/`).
191
- 3. If the resulting list is empty, emit "diff mode: no changed files vs. main; rerun without `diff` or specify a commit range" per the error table and exit before dispatch.
192
- 4. Otherwise, this list is passed to each subagent as the `Scope` (the `Diff:` form of the prompt template in Phase 4).
228
+ 1. Resolve the base branch: `git symbolic-ref --short refs/remotes/origin/HEAD` (strip the `origin/` prefix); if that fails, probe `git rev-parse --verify main` then `master` and use the first that exists. Run `git diff --name-only origin/<base>...HEAD` (fall back to `git diff --name-only <base>...HEAD` if no `origin` remote).
229
+ 2. If git itself errors (not a repo, unknown revision), report the git error verbatim and exit; do not conflate this with the empty-diff case below.
230
+ 3. Filter out paths that no longer exist (deleted files) and paths in the standard ignored set (`node_modules/`, `.git/`, `dist/`, `build/`, `coverage/`, `tests/fixtures/`). If `scope.module` is also set, keep only paths under the module path (diff and module compose as an intersection).
231
+ 4. If the resulting list is empty, emit "diff mode: no changed files vs. <base><if module: within <module-path>>; rerun without `diff` or specify a commit range" per the error table and exit before dispatch.
232
+ 5. Otherwise, this list is passed to each subagent as the `Scope` (the `Diff:` form of the prompt template in Phase 4).
193
233
 
194
234
  ## Phase 4: Dispatch subagents
195
235
 
@@ -202,14 +242,15 @@ You are running the <audit-id> audit on a codebase.
202
242
 
203
243
  # Principles to apply
204
244
 
205
- <paste the full content of skills/<audit-id>/SKILL.md verbatim>
245
+ <paste the full content of <skills-root>/<audit-id>/SKILL.md verbatim (see Phase 3, Locating audit content)>
206
246
 
207
247
  # Scope
208
248
 
209
249
  <one of:>
210
250
  - Whole repo at <repo-root-absolute-path>.
211
251
  - Module: <module-path-absolute>.
212
- - Diff: only the following files changed vs. main: <newline-separated list>.
252
+ - Diff: only the following files changed vs. <base>: <newline-separated list>.
253
+ - Diff within module <module-path>: only the following changed files under it: <newline-separated list>.
213
254
 
214
255
  # Method
215
256
 
@@ -245,9 +286,9 @@ Schema:
245
286
  - Do not include findings outside the declared scope.
246
287
  ````
247
288
 
248
- **Concurrency**: dispatch all subagents in a single tool-call batch. Subagents do not communicate with each other.
289
+ **Concurrency**: dispatch all subagents in a single tool-call batch. Subagents do not communicate with each other. (On a harness without a parallel subagent tool, run the same prompts sequentially and collect outputs identically.)
249
290
 
250
- **Soft cap**: if a subagent has not returned within 5 minutes, drop it. The orchestrator does not have a wall-clock timer; this cap is enforced by treating any subagent that fails to return cleanly as `status: failed` and continuing.
291
+ **Stuck subagents**: treat any subagent that fails to return cleanly as `status: failed` and continue; do not wait on or retry a hung dispatch.
251
292
 
252
293
  **Retry policy:**
253
294
  - If a subagent returns prose, markdown fences, or text that does not parse as JSON, retry **once** with this follow-up prompt:
@@ -266,14 +307,14 @@ Schema:
266
307
  ```json
267
308
  {
268
309
  "audit": "<audit-id>",
269
- "status": "ok" | "failed",
310
+ "status": "ok" | "failed" | "not-installed",
270
311
  "reason": "<string or null>",
271
312
  "dropped_findings": <int, count of findings dropped at validation>,
272
313
  "retried": <bool, true if recovered on attempt 2>
273
314
  }
274
315
  ```
275
316
 
276
- One such record is emitted per dispatched audit. Successful runs use `status: "ok"`, `reason: null`. Failed runs use `status: "failed"` and a reason string (e.g. `"timeout"`, `"crashed"`, `"malformed JSON after 2 attempts"`).
317
+ One such record is emitted per dispatched audit (plus one per routed-but-not-installed audit from Phase 3). Successful runs use `status: "ok"`, `reason: null`. Failed runs use `status: "failed"` and a reason string (e.g. `"timeout"`, `"crashed"`, `"malformed JSON after 2 attempts"`).
277
318
 
278
319
  ## Phase 5: Synthesize
279
320
 
@@ -293,11 +334,11 @@ Merge all `Finding[]` arrays into a single `Report` object.
293
334
  - `symbol_boost`: if the finding has a non-null `symbol` AND at least one other finding on the same file shares that symbol AND comes from a different audit, multiply by `1.25`.
294
335
  - `weight = severity_weight * convergence_weight * symbol_boost`.
295
336
  4. **Per-file score**: sum of `weight` across all findings on that file.
296
- 5. **Sort files** by per-file score descending. For each file in the result, attach `top_issue` = the `principle` of the highest-weight finding on that file. Take the top N files (`N = 5` if `mode === "quick"`, else `10`; or all if fewer) for the "files most worth fixing" section.
297
- 6. **Sort findings** flat by `weight` descending. For each finding in the result, attach `convergence` = `"<distinct_audits_on_file> audits on this file"` (e.g. "4 audits on this file"). Take the top M findings (`M = 5` if `mode === "quick"`, else `25`; or all if fewer) for the "top cross-cutting findings" section.
337
+ 5. **Sort files** by per-file score descending; break ties by file path ascending (deterministic output across runs). For each file in the result, attach `top_issue` = the `principle` of the highest-weight finding on that file. Take the top N files (`N = 5` if `mode === "quick"`, else `10`; or all if fewer) for the "files most worth fixing" section.
338
+ 6. **Sort findings** flat by `weight` descending; break ties by severity (high > med > low), then file path ascending, then line ascending (nulls last). For each finding in the result, attach `convergence` = `"<distinct_audits_on_file> audits on this file"` (e.g. "4 audits on this file"). Take the top M findings (`M = 5` if `mode === "quick"`, else `25`; or all if fewer) for the "top cross-cutting findings" section.
298
339
  7. **Build per-audit summaries**: for each audit ID present in the input map and with `findings.length > 0`, compute `{finding_count: <int>, unique_file_count: <int>, findings_by_file: [{file, findings: [...sorted by weight desc]}, ...]}`. Audits that returned `[]` are omitted from `by_audit_summary` (so Phase 6's per-axis index does not link to empty files).
299
340
  8. **Preserve raw findings**: copy the input `audit-id -> Finding[]` map to `Report.by_audit` verbatim (audits with empty arrays are still present here for completeness; the per-axis index in Phase 6 reads `by_audit_summary` instead, which omits empty audits).
300
- 9. **Compute `audits_na`**: `audits_na = catalogue_audit_count - audits_dispatched_count` where `catalogue_audit_count` is the total number of audits in the catalogue with `applies` and `quick` fields, and `audits_dispatched_count` is the number routed by Phase 3 (regardless of pass/fail).
341
+ 9. **Compute `audits_na`**: `audits_na = catalogue_audit_count - audits_dispatched_count`, where `catalogue_audit_count` is 26 (the routing-table total) and `audits_dispatched_count` is the number actually dispatched in Phase 4 (regardless of pass/fail). Routed-but-not-installed audits are never dispatched, so they count toward `audits_na`; also list them in `failures` with reason `"not installed"`. Invariant: `audits_succeeded + audits_failed + audits_na = audits_total = catalogue_audit_count`.
301
342
 
302
343
  **Output `Report` structure:**
303
344
 
@@ -309,8 +350,8 @@ Merge all `Finding[]` arrays into a single `Report` object.
309
350
  "scope": { "module": "...", "focus": "..." },
310
351
  "audits_succeeded": 14,
311
352
  "audits_failed": 2,
312
- "audits_na": 4,
313
- "audits_total": 30,
353
+ "audits_na": 10,
354
+ "audits_total": 26,
314
355
  "failures": [{"audit": "event-design-reviewer", "reason": "malformed JSON after 2 attempts"}],
315
356
  "retries": [{"audit": "rams-design-audit", "recovered_on_attempt": 2}],
316
357
  "findings_high": 47,
@@ -374,7 +415,7 @@ The only filesystem-mutating phase. Produces one top-level rollup and one per-ax
374
415
  - `docs/improvements/YYYY-MM-DD/audit.md` (top-level rollup).
375
416
  - `docs/improvements/YYYY-MM-DD/audit/<audit-id>.md` (one per audit that returned at least one finding).
376
417
 
377
- If `docs/improvements/YYYY-MM-DD/` already exists for today, append a numeric suffix: `audit-2.md`, `audit-3.md`, etc. Do not overwrite a previous run.
418
+ If today's directory already contains a run, suffix BOTH outputs so nothing from the earlier run is overwritten: `audit-2.md` with per-axis files under `audit-2/<audit-id>.md` (then `audit-3.md` + `audit-3/`, etc.). Links inside each rollup point at its own per-axis directory.
378
419
 
379
420
  **`audit.md` template:**
380
421
 
@@ -491,11 +532,11 @@ printTerminalSummary(report)
491
532
  When `parsed.mode === "interactive"`:
492
533
 
493
534
  1. Show the user the detected stack and the audit set the orchestrator would run by default.
494
- 2. Ask, one at a time:
535
+ 2. Ask, one at a time (any `focus`/`module` provided on the command line is presented as the default answer):
495
536
  - "Anything you've been losing time on lately?" (free text, suggests a `focus` area)
496
537
  - "Any directory you want to scope to?" (path or skip; sets `module`)
497
538
  - "Quick scan or thorough?" (sets `quick` or full)
498
- 3. Apply the answers as if they were args, re-route, then proceed to dispatch.
539
+ 3. Apply the answers as if they were args (scope values are honored in every mode), re-route, then proceed to dispatch.
499
540
  4. Confirm the final audit list before dispatching: "Running these audits: <list>. OK?"
500
541
 
501
542
  Interactive mode is the only path that reroutes after Phase 3.
@@ -513,4 +554,4 @@ Interactive mode is the only path that reroutes after Phase 3.
513
554
  | Subagent malformed JSON | 4 | Retry up to 1 time (2 attempts total). Then drop. |
514
555
  | Finding fails schema | 5 | Drop bad finding only. Note count in metadata. |
515
556
  | Cannot write report file | 6 | Fall back to terminal-only. Print full report. Surface write error. |
516
- | `index.json` missing/malformed | 3 | Hard fail. Suggest reinstalling. |
557
+ | Routed audit's SKILL.md missing on disk | 3 | Skip that audit, record `{audit, status: "not-installed"}`, suggest installing it. Continue; hard fail only if no routed audit is installed. |
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "improve-my-codebase",
3
- "version": "1.0.1",
3
+ "version": "1.1.0",
4
4
  "description": "Orchestrator skill that runs every applicable audit skill in parallel and produces a prioritized, convergence-ranked improvement report",
5
5
  "author": "iceinvein",
6
6
  "type": "prompt",