ts-reviewer 3.1.0 → 3.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,13 +6,26 @@ Built for one fixed stack — **TypeScript 5.9.x, ES2024, Node 24** — without
6
6
 
7
7
  ## What It Does
8
8
 
9
- Three modes, one skill:
9
+ Four modes, one skill:
10
10
 
11
11
  | Mode | What happens |
12
12
  |---|---|
13
13
  | **scan** | Analyzes the codebase and writes a prioritized report to `code-smells/report.md` |
14
+ | **investigate** | Reads the report and decides, from tests, git history, decision records, and callers, whether each finding is a real defect or deliberate. Changes no code |
14
15
  | **fix** | Reads the report and applies fixes file-by-file with tsc/lint/test verification |
15
- | **auto** | Runs scan, asks you to confirm, fixes everything, deletes the report if clean |
16
+ | **auto** | Runs scan, investigates, asks you to confirm, fixes everything, deletes the report if clean |
17
+
18
+ ## What's New
19
+
20
+ **3.5.0 — pick what runs, and sturdier large runs.** `--domains security,boundary-validation` runs only those domains, and `--pick` asks you in a multi-select. The skill lint drops findings outside the pick, and skips itself when no picked domain owns a lint line. A new module, such as a future framework checklist, joins the menu through its `pass_groups` row. From a run on a 98-file monorepo: a database row typed by a generic and an untyped `JSON.parse` now grade High, a Next.js package is out of scope, `tools/check-passes.mjs` repairs and checks the pass files before the merge, every pass gets a fresh agent, and every Recurring Pattern row lists its sites. The main agent now writes 1 pass plan and its decisions, and 2 tools write the pass prompts and the report: on the recall corpus the main agent spends 35% less, on the monorepo 37% less, and every report validates on the first try.
21
+
22
+ **3.4.0 — a cheaper scan.** A pinned ESLint + typescript-eslint config now runs inside the scan (1 approval, through `npx`), and the 49 checklist lines it decides by rule leave the AI passes. A default scan runs 5 pass groups instead of 9 domain passes, the async and error group reads only the files that can hold its patterns, and the config and dependency group reads no `.ts` file. The pass model is yours to choose once per project. The report format is unchanged. Measured on a 98-file monorepo, the scan spent 8.6% less and finished 10 minutes sooner with the skill lint. On the recall corpus in `fixtures/recall-corpus/`, Sonnet passes halved what the passes cost and kept every High and Highest finding.
23
+
24
+ **3.3.0 — investigate mode.** Before a fix changes flagged code, the skill now checks whether the pattern is there on purpose: a test that pins it, a commit message that explains it, an ADR that decides it. Deliberate code is left alone and gets a `// Deliberate:` comment citing the evidence, so the next scan does not flag it again. See [Investigate](#investigate--why-is-this-code-this-way).
25
+
26
+ **3.2.0 — hot paths.** Mark performance-critical code with `/** @hotpath */`. A fix that would add an allocation, a validation, or an extra pass there is redesigned to pay that cost outside the hot path, or handed to you as a choice when it cannot be. See [Hot paths](#hot-paths--hotpath).
27
+
28
+ Both are optional: a project with no `@hotpath` markers, no tests, and no history gets today's behaviour plus 1 verdict line per finding.
16
29
 
17
30
  The review covers nine domains by default, each with its own detailed checklist. Add `--arch` or `--full` to include architecture analysis:
18
31
 
@@ -67,6 +80,21 @@ You can still copy the `ts-reviewer/` folder directly into the skill directory f
67
80
 
68
81
  ## Usage
69
82
 
83
+ The usual workflow is 3 requests in 1 session, or 1 request in auto mode:
84
+
85
+ ```
86
+ Review my TypeScript code → code-smells/report.md
87
+ Investigate the report → 1 Verdict line per finding, no code change
88
+ Fix the report → fixes, comments on deliberate code, audit trail
89
+ ```
90
+ ```
91
+ Review and fix my TypeScript code → all of the above, with 1 confirmation before the fix
92
+ ```
93
+
94
+ Read the report between the steps: it is the work plan, and you can delete findings or edit a verdict before fix runs. Investigate is optional — `Fix the report` straight after a scan works as in earlier versions.
95
+
96
+ You do not start any sub-agents yourself. The scan launches its own analysis passes as sub-agents (`--agents N`, default 3); investigate and fix run in the main agent.
97
+
70
98
  ### Scan — find issues
71
99
 
72
100
  Just ask Claude to review your code:
@@ -86,6 +114,12 @@ With Architecture active, the same directory also holds project discovery, Knip,
86
114
 
87
115
  The analysis passes run in waves. `--agents N` sets how many run at once (default 3; `--agents 1` runs them one at a time in the main agent). Each pass writes its own findings file under `code-smells/passes/`, and the queue in `code-smells/passes/queue.md` tracks which passes are done, so an interrupted scan does not lose finished work.
88
116
 
117
+ The scan also runs a skill lint: a pinned ESLint + typescript-eslint config (`ts-reviewer/tools/eslint.config.mjs`), through `npx -y eslint@10 typescript-eslint@8 typescript@5.9`, after 1 approval. The checklist lines it decides by rule are marked `lint-owned` in the references, its findings land in the report like any pass, and the AI passes skip those lines. Declined or failed, the scan falls back to the passes for every line. Your project's own linter still runs as before.
118
+
119
+ The main agent writes its decisions, not the report: `tools/pass-prompts.mjs` fills the pass prompts from 1 plan, and `tools/build-report.mjs` applies the deduplication, merge, Recurring Pattern, and sorting steps to the pass files and renders the report, which `validate-report.mjs` then checks.
120
+
121
+ The passes run in groups: Type Safety with Boundary Validation, Async Patterns with Error Handling, Config with Dependency Hygiene, Modernization with Code Quality, and Security and Architecture alone.
122
+
89
123
  #### Domain flags
90
124
 
91
125
  By default, only the nine core domains run. Use flags to control which domains are active:
@@ -96,6 +130,8 @@ By default, only the nine core domains run. Use flags to control which domains a
96
130
  | `--arch` | Architecture only (shallow modules, coupling, dependency direction, seams) |
97
131
  | `--full` | All ten domains |
98
132
  | `--no-arch` | The nine core domains — overrides `--arch`, `--full`, and any phrase that would enable architecture |
133
+ | `--domains <slugs>` | Only the named domains, by slug (`security`, `type-safety`, `boundary-validation`, ...) or by pass group (`type-safety+boundary-validation`). `--no-arch` still removes Architecture |
134
+ | `--pick` | Asks which pass groups to run, in a multi-select (a numbered list on Codex). No answer runs the default set |
99
135
 
100
136
  Examples:
101
137
 
@@ -123,7 +159,7 @@ Apply fixes from code-smells/report.md
123
159
  The fix workflow:
124
160
  1. Parses the report as a work plan
125
161
  2. Runs existing tests to capture a baseline (knows what was already failing)
126
- 3. Fixes issues file-by-file, writes regression tests, runs `tsc` after each file
162
+ 3. Fixes issues file-by-file, writes regression tests, runs `tsc` after each file. A finding investigate called deliberate gets at most a comment; a fix on a hot path is redesigned first (see [Hot paths](#hot-paths--hotpath))
127
163
  4. Runs linter, fixes lint errors
128
164
  5. Runs full test suite, compares with baseline, fixes any regressions it caused
129
165
  6. Repeats verification up to 5 iterations
@@ -140,7 +176,43 @@ Review and fix my TypeScript code
140
176
  Auto-fix code smells
141
177
  ```
142
178
 
143
- Runs scan, shows you the summary, asks if you want to proceed with fixes, then runs the full fix cycle. If everything is clean afterward, the report is deleted.
179
+ Runs scan, investigates every finding, shows you the summary, asks if you want to proceed with fixes, then runs the full fix cycle. If everything is clean afterward, the report is deleted.
180
+
181
+ ### Investigate — why is this code this way?
182
+
183
+ ```
184
+ Investigate the report
185
+ ```
186
+ ```
187
+ Investigate the ts-reviewer report: which findings are deliberate?
188
+ ```
189
+
190
+ Run it after a scan, before fix. It needs `code-smells/report.md`, and it works best in a git repository with real commit messages and a test suite — those are its evidence.
191
+
192
+ Before a fix changes flagged code, investigate asks whether the pattern is there on purpose. For each finding in `code-smells/report.md` it reads 5 sources, cheapest first, and stops at the first that names the flagged behaviour: a comment at the site, a test that calls the function, the commit that introduced the exact lines (`git log -L`), a decision record (`docs/adr/`, `ARCHITECTURE.md`, ...), and the callers. It writes 1 `**Verdict:**` line per finding with a pointer to that source, and changes no code.
193
+
194
+ | Verdict | Decided by | What fix mode does |
195
+ |---|---|---|
196
+ | `defect` | a test, a record, or a caller shows the failure | fixes it as reported |
197
+ | `deliberate-recorded` | a decision record names the behaviour as wanted | no change, `[SKIPPED: deliberate, <pointer>]` |
198
+ | `deliberate-unrecorded` | a test or a commit message names it as wanted, and nothing at the site does | adds `// Deliberate: <behaviour>. Evidence: <pointer>.` above the line, no other change |
199
+ | `unreachable` | every caller is known, and none reaches the failure | fixes it when the fix is free; on a hot path a costly fix is skipped without asking |
200
+ | `unknown` | no source decides | fixes it as today; the verdict tells you the fix rests on no evidence |
201
+
202
+ A commit message counts only when it names the behaviour ("wip" does not), and a test is read before the history, so a test that asserts the opposite wins. The comment a `deliberate-unrecorded` verdict leaves is what makes the next scan drop the finding.
203
+
204
+ ### Hot paths — `@hotpath`
205
+
206
+ Some correct fixes are the wrong default in a frame loop or a message handler: `toSorted()` allocates on every call, a schema parse validates every message. Mark such code with a JSDoc tag:
207
+
208
+ ```typescript
209
+ /** @hotpath */
210
+ export function updateFrame(entities: Entity[], dt: number): void { /* ... */ }
211
+ ```
212
+
213
+ The tag is optional. `@hotpath` on a function, method, or class covers its body; in a file's leading comment it covers the whole file. Without any marker, only loop bodies and iteration callbacks (`map`, `forEach`, `sort`, ...) count as hot. A project running `eslint-plugin-jsdoc` with `check-tag-names` has to declare `hotpath` in `definedTags`.
214
+
215
+ A finding on a hot path whose fix adds a per-call cost carries a `**Hot path:**` line in the report. Fix mode does not apply such a fix as written: it walks a ladder of 7 rungs — remove the case through types, move the cost to the boundary, hoist it, reuse a module-owned buffer, check it in development builds only, split off the common case — and applies the first that closes the finding. When none does (rung 7), the code stays untouched and you are shown both variants: the reported fix with its cost, and the current code with its defect. You pick one; with no answer the entry is recorded as `[SKIPPED: rung 7 <kind>]`. A `bench` or `benchmark` script in `package.json` is run before and after, and both numbers go on every designed fix.
144
216
 
145
217
  ## Scope Modes
146
218
 
@@ -199,11 +271,15 @@ cnlp/ # the CNL-P format the skill files are wri
199
271
 
200
272
  ts-reviewer/
201
273
  ├── SKILL.md # Main skill file — mode routing, workflow orchestration
202
- ├── tools/ # Mechanical pre-pass and report validator — plain Node, no dependencies
274
+ ├── tools/ # Mechanical steps of the scan — plain Node, no dependencies
203
275
  │ ├── discover-projects.mjs # Finds the TypeScript projects and their source roots
204
276
  │ ├── co-change.mjs # Git co-change pairs across directory boundaries
205
277
  │ ├── run-cruise.mjs # dependency-cruiser graphs, metrics, and Mermaid diagrams per project
206
278
  │ ├── classify-run.mjs # Reads a tool run by its output, not its exit code
279
+ │ ├── eslint.config.mjs, lint-rules.mjs, lint-pass.mjs # The skill lint and its pass file
280
+ │ ├── pass-prompts.mjs # Writes the pass queue and 1 filled prompt per pass
281
+ │ ├── check-passes.mjs # Repairs and checks the pass files before the merge
282
+ │ ├── build-report.mjs # Applies the merge steps and renders code-smells/report.md
207
283
  │ └── validate-report.mjs # Checks code-smells/report.md against the report contract
208
284
  └── references/
209
285
  ├── type-safety.md # Checklist: any, unknown, casts, !, exhaustiveness, branded types
@@ -216,10 +292,18 @@ ts-reviewer/
216
292
  ├── tsconfig.md # Checklist: strict flags, target/lib, module resolution, deprecated
217
293
  ├── dependency-hygiene.md # Checklist: lockfiles, versions, npm audit, dependency choice
218
294
  ├── architecture.md # Checklist: shallow modules, coupling, dependency direction, seams
219
- └── fix-workflow.md # Complete fix protocol: tests, verification, rollback
295
+ ├── fix-workflow.md # Complete fix protocol: tests, verification, rollback
296
+ ├── fix-design.md # Stack-free ladder for designing a fix on a hot path
297
+ ├── stack-cost.md # The @hotpath marker, cost kinds, rung forms, bench command, evidence sources
298
+ └── investigate.md # Stack-free verdicts: why flagged code is the way it is
299
+
300
+ docs/ # design proposals behind each feature, with their decisions
301
+ fixtures/ # throwaway projects + answer keys the features were tested against (unpublished)
302
+ ├── cost-corpus/ # → hot paths, 3.2.0
303
+ └── intent-corpus/ # → investigate, 3.3.0
220
304
  ```
221
305
 
222
- **SKILL.md** is the orchestrator — it routes between scan/fix/auto modes, detects domain flags (`--arch`, `--full`), defines scope detection, severity scale, and report format.
306
+ **SKILL.md** is the orchestrator — it routes between scan/investigate/fix/auto modes, detects domain flags (`--arch`, `--full`), defines scope detection, severity scale, and report format.
223
307
 
224
308
  **Reference files** contain the detailed checklists and protocols. Each analysis agent reads only the reference file relevant to its domain, keeping context focused. Architecture analysis is opt-in and loaded only when the domain is active.
225
309
 
@@ -247,25 +331,34 @@ The test catches a check line that lost its severity, a block the profile does n
247
331
 
248
332
  ### Scan mode
249
333
 
250
- 1. **Discovery** — detects domain flags, maps the project, reads tsconfig.json, detects linter and test runner, and asks once before downloading a missing architecture tool.
251
- 2. **Diagnostics** — runs `tsc --noEmit`, linter, and LSP diagnostics (if available); compiler and linter output is cached under `code-smells/passes/` and reused on a resume of the same commit.
334
+ 1. **Discovery** — detects domain flags, maps the project, reads tsconfig.json, detects linter and test runner, asks the pass model once per project, and asks once before downloading the skill lint or a missing architecture tool.
335
+ 2. **Diagnostics** — runs `tsc --noEmit`, the project linter, the skill lint, and LSP diagnostics (if available). The skill lint's findings become the pass `lint-skill`; compiler and linter output is cached under `code-smells/passes/` and reused on a resume of the same commit.
252
336
  3. **Architecture pre-pass** — when active, writes bounded Knip, graph, metric, co-change, rule, and Mermaid artifacts under `code-smells/`, with project coverage and bounded failure diagnostics.
253
- 4. **Analysis** — specialized passes judge the candidates against the active checklists, running in waves of `--agents` at a time; each pass writes its own `code-smells/passes/<id>.jsonl`, and `passes/queue.md` marks which are done, so a stopped run resumes from the last checkpoint. Tool output is never a finding by itself.
337
+ 4. **Analysis** — specialized passes, 1 per group of domains, judge the candidates against the active checklists, skipping the lines the skill lint owns, running in waves of `--agents` at a time; each pass writes its own `code-smells/passes/<id>.jsonl`, and `passes/queue.md` marks which are done, so a stopped run resumes from the last checkpoint. Tool output is never a finding by itself.
254
338
  5. **Report** — deduplicates, applies severity boost (scoped modes), consolidates recurring patterns, enforces a noise budget, writes `code-smells/report.md`, and validates its contract before the scan succeeds. Architecture findings appear in a separate `## Architecture Opportunities` section at the end.
255
339
 
256
- Validate a report directly with `node ts-reviewer/tools/validate-report.mjs --repo . --report code-smells/report.md`. It checks headings, counts, finding anchors, architecture fields, and linked artifacts without adding a dependency. An **error** is a defect of the report that rewriting it fixes; a **warning** names an outcome of the mechanical pre-pass — a graph with no diagram, say — that the report cannot fix, and warnings do not fail the run.
340
+ Validate a report directly with `node ts-reviewer/tools/validate-report.mjs --repo . --report code-smells/report.md`. It checks headings, counts, finding anchors, architecture fields, and linked artifacts without adding a dependency, and it reads both the scan report and the audit trail a fix run leaves in its place. An **error** is a defect of the report that rewriting it fixes; a **warning** names an outcome of the mechanical pre-pass — a graph with no diagram, say — that the report cannot fix, and warnings do not fail the run.
341
+
342
+ ### Investigate mode
343
+
344
+ 1. Reads `references/investigate.md` and the evidence locations in `references/stack-cost.md`
345
+ 2. For each `###` finding, reads the sources in order — site comment, tests, `git log -L` on the exact lines, decision records, callers — and stops at the first that names the flagged behaviour
346
+ 3. Writes `**Verdict:** <verdict> | **Evidence:** <source> <pointer>` into the entry, and validates the report
347
+ 4. Changes no source file: `git diff` is the same before and after
257
348
 
258
349
  ### Fix mode
259
350
 
260
351
  1. Validates `code-smells/report.md` and stops before changing code when the report is invalid
261
352
  2. Parses the report as the work plan
262
- 3. Captures test baseline (runs tests before changes)
263
- 4. Applies fixes bottom-to-top within each file (so line numbers don't shift)
264
- 5. Writes regression tests for each testable fix
265
- 6. Runs `tsc --noEmit` after each file
266
- 7. Runs the full verification loop: tsc + linter + test suite (max 5 iterations)
267
- 8. Compares test results with baseline — only fixes regressions it caused
268
- 9. Updates or deletes the report, keeps the remaining `code-smells/` artifacts, and asks before removing them
353
+ 3. Captures test baseline (runs tests before changes), and the `bench` script when a finding is on a hot path
354
+ 4. Closes deliberate findings first: no change for `deliberate-recorded`, 1 `// Deliberate:` comment for `deliberate-unrecorded`
355
+ 5. Applies fixes bottom-to-top within each file (so line numbers don't shift); a fix on a hot path goes through the 7-rung design first
356
+ 6. Writes regression tests for each testable fix
357
+ 7. Runs `tsc --noEmit` after each file
358
+ 8. Runs the full verification loop: tsc + linter + test suite (max 5 iterations)
359
+ 9. Compares test results with baseline — only fixes regressions it caused
360
+ 10. Shows you the rung 7 choices, runs the `bench` script again, and writes both numbers on every designed fix
361
+ 11. Updates or deletes the report, keeps the remaining `code-smells/` artifacts, and asks before removing them
269
362
 
270
363
  ## Tips
271
364
 
@@ -275,9 +368,15 @@ Validate a report directly with `node ts-reviewer/tools/validate-report.mjs --re
275
368
 
276
369
  - **Claude Code users** — `CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS` in `settings.json` under `env` caps sub-agents for every session on the host. It is independent of `--agents`, which caps one review run and works in every supported agent.
277
370
 
371
+ - **Pick the pass model once** — the first scan asks which model and effort the analysis passes use, and writes the answer to `.claude/agents/ts-reviewer-scout.md` (Claude Code: `model`, `effort`) or `.codex/agents/ts-reviewer-scout.toml` (Codex: `model`, `model_reasoning_effort`). A smaller model there, for example Sonnet at `high`, costs less, while the main agent keeps verifying every finding. Delete the file to be asked again, or pass `--scout <model>` for 1 run. Architecture always runs on the main agent's model.
372
+
278
373
  - **Commit before running fix** — so you can `git diff` to review changes and `git checkout -- .` to revert if needed.
279
374
 
280
- - **Edit the report before fix** — since fix uses `code-smells/report.md` as its work plan, you can delete issues you don't want fixed, change severities, or add notes before running fix.
375
+ - **Edit the report before fix** — since fix uses `code-smells/report.md` as its work plan, you can delete issues you don't want fixed, change severities, change a `Verdict` line, or add notes before running fix.
376
+
377
+ - **Leave evidence of intent** — a test that asserts the behaviour, a commit message that names it, or an ADR in `docs/adr/` is what investigate reads. A code comment at the site is the strongest: the scan drops the finding outright.
378
+
379
+ - **Mark hot paths once** — `/** @hotpath */` on a frame loop, parser, or message handler keeps every future fix there allocation-aware. Unmarked loops are still treated as possibly hot.
281
380
 
282
381
  - **Scoped review for PRs** — `"review my branch against main"` is the most practical mode for day-to-day use. Full codebase audits are better suited for periodic health checks.
283
382
 
@@ -286,6 +385,7 @@ Validate a report directly with `node ts-reviewer/tools/validate-report.mjs --re
286
385
  - TypeScript 5.9.x project targeting ES2024 on Node 24
287
386
  - Git repository (for scoped modes and safe revert during fix)
288
387
  - Node 24 with `npx` available (for tsc, linter)
388
+ - Optional: a `bench` or `benchmark` script in `package.json`, for before/after numbers on hot-path fixes
289
389
  - Claude Code (recommended) or any Claude interface with skill support
290
390
 
291
391
  ## License
package/package.json CHANGED
@@ -1,12 +1,12 @@
1
1
  {
2
2
  "name": "ts-reviewer",
3
- "version": "3.1.0",
3
+ "version": "3.5.0",
4
4
  "description": "Install the TypeScript Code Reviewer skill for Claude Code, Codex, or Antigravity",
5
5
  "license": "MIT",
6
6
  "type": "module",
7
7
  "repository": {
8
8
  "type": "git",
9
- "url": "https://github.com/VirtualMaestro/ts-reviewer"
9
+ "url": "git+https://github.com/VirtualMaestro/pure-typescript-reviewer.git"
10
10
  },
11
11
  "bin": {
12
12
  "ts-reviewer": "dist/cli.js"