any-doctor 0.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (74) hide show
  1. package/CONTEXT.md +128 -0
  2. package/README.md +68 -0
  3. package/bin/capabilities.d.ts +15 -0
  4. package/bin/capabilities.js +131 -0
  5. package/bin/cli.d.ts +2 -0
  6. package/bin/cli.js +426 -0
  7. package/bin/clipboard.d.ts +1 -0
  8. package/bin/clipboard.js +10 -0
  9. package/bin/contract.d.ts +134 -0
  10. package/bin/contract.js +70 -0
  11. package/bin/dashboard.d.ts +108 -0
  12. package/bin/dashboard.js +718 -0
  13. package/bin/discover.d.ts +24 -0
  14. package/bin/discover.js +87 -0
  15. package/bin/doctor-loader.d.mts +1 -0
  16. package/bin/doctor-loader.mjs +161 -0
  17. package/bin/engine.d.ts +18 -0
  18. package/bin/engine.js +22 -0
  19. package/bin/fuzzy.d.ts +2 -0
  20. package/bin/fuzzy.js +31 -0
  21. package/bin/import-guard.mjs +31 -0
  22. package/bin/keys.d.ts +2 -0
  23. package/bin/keys.js +72 -0
  24. package/bin/palette.d.ts +7 -0
  25. package/bin/palette.js +14 -0
  26. package/bin/picker.d.ts +12 -0
  27. package/bin/picker.js +82 -0
  28. package/bin/report.d.ts +18 -0
  29. package/bin/report.js +159 -0
  30. package/bin/runner.d.ts +58 -0
  31. package/bin/runner.js +271 -0
  32. package/bin/score.d.ts +13 -0
  33. package/bin/score.js +39 -0
  34. package/bin/sdk.d.ts +5 -0
  35. package/bin/sdk.js +95 -0
  36. package/bin/search-host.d.ts +6 -0
  37. package/bin/search-host.js +56 -0
  38. package/bin/select.d.ts +35 -0
  39. package/bin/select.js +45 -0
  40. package/bin/tty.d.ts +38 -0
  41. package/bin/tty.js +94 -0
  42. package/docs/REPAIR-LOG.md +45 -0
  43. package/docs/RESULTS.md +70 -0
  44. package/docs/decisions.md +450 -0
  45. package/docs/example-catalog.md +122 -0
  46. package/docs/features.md +67 -0
  47. package/docs/first-shot-results.md +18 -0
  48. package/docs/intents.md +21 -0
  49. package/docs/kill-test.md +54 -0
  50. package/docs/research.md +66 -0
  51. package/docs/vision.md +83 -0
  52. package/doctors/AGENTS.md +103 -0
  53. package/doctors/api-route-files-do-import.fixtures.mjs +61 -0
  54. package/doctors/api-route-files-do-import.mjs +26 -0
  55. package/doctors/async-doctor.fixtures.mjs +147 -0
  56. package/doctors/async-doctor.mjs +295 -0
  57. package/doctors/convex-doctor.fixtures.mjs +177 -0
  58. package/doctors/convex-doctor.mjs +223 -0
  59. package/doctors/date-now-used-inside-effect.fixtures.mjs +46 -0
  60. package/doctors/date-now-used-inside-effect.mjs +132 -0
  61. package/doctors/json-parse-calls-llm-api.fixtures.mjs +37 -0
  62. package/doctors/json-parse-calls-llm-api.mjs +85 -0
  63. package/doctors/route-handlers-touch-database-before.fixtures.mjs +58 -0
  64. package/doctors/route-handlers-touch-database-before.mjs +98 -0
  65. package/doctors/z-record-called-with-single.fixtures.mjs +28 -0
  66. package/doctors/z-record-called-with-single.mjs +19 -0
  67. package/fixtures/sample-app/src/hooks/useChat.ts +15 -0
  68. package/fixtures/sample-app/src/lib/ai/client.ts +5 -0
  69. package/fixtures/sample-app/src/schemas/user.ts +6 -0
  70. package/fixtures/sample-app/src/services/chat.ts +17 -0
  71. package/fixtures/sample-app/src/services/user.ts +10 -0
  72. package/fixtures/sample-app/src/utils/sync.ts +16 -0
  73. package/package.json +42 -0
  74. package/skill/any-doctor.skill.md +188 -0
@@ -0,0 +1,450 @@
1
+ # Decision log
2
+
3
+ Append-only. Each entry: context → decision → consequences. New sessions
4
+ should read this file first and *not* relitigate closed decisions.
5
+
6
+ ---
7
+
8
+ ## D1 — Open source tool, not a startup
9
+
10
+ **Date:** 2026-08-28
11
+
12
+ **Context:** Idea emerged from studying React Doctor. Explicitly not pursuing
13
+ as a venture.
14
+
15
+ **Decision:** Build as an OSS project. Optimize for usefulness, trust, and a
16
+ compelling demo — not defensibility. Crowded space is validation, not threat.
17
+
18
+ **Consequences:** No moat engineering. Prioritize boring stable artifacts over
19
+ platform ambitions. The launch content (measured precision/recall) doubles as
20
+ marketing.
21
+
22
+ ---
23
+
24
+ ## D2 — Agent-native architecture: harness + skill, no shipped LLM
25
+
26
+ **Date:** 2026-08-28
27
+
28
+ **Context:** A standalone CLI calling an LLM API means key management, cost,
29
+ model churn, and we'd be maintaining an AI pipeline nights-and-weekends. The
30
+ alternative: the user's own agent does generation.
31
+
32
+ **Decision:** The OSS artifact is a rule format + fixture harness + skill file
33
+ (+ thin CLI). Agents (Claude Code, Cursor, Codex) generate rules into the
34
+ format. CI runs saved rules with zero inference.
35
+
36
+ **Consequences:** Small stable core to maintain. Generation quality depends on
37
+ the user's agent, which keeps improving for free. v1 could literally be a
38
+ skill + harness. Downside: generation UX varies by agent; acceptable for OSS.
39
+
40
+ ---
41
+
42
+ ## D3 — ast-grep as v1 executor; YAML artifacts
43
+
44
+ **Date:** 2026-08-28
45
+
46
+ **Context:** Chosen over building on oxc/oxlint. Three sub-decisions were
47
+ conflated and separated: (1) artifact format the LLM emits, (2) analysis power
48
+ needed, (3) API stability.
49
+
50
+ **Decision:** Generated rules are **ast-grep YAML** (declarative: pattern +
51
+ constraints + relational ops), with `ast-grep test` snapshot fixtures. JS
52
+ visitors are an escape hatch, not the default.
53
+
54
+ **Rationale:** Constrained artifacts generate more reliably (fewer degrees of
55
+ freedom = fewer ways to be subtly wrong); auditable at a glance in a PR;
56
+ data-can't-execute (a generated YAML file can't phone home — a generated JS
57
+ plugin is a small supply-chain incident in a stranger's repo); native test
58
+ harness gives us the fixture story free; stable format for years vs oxlint's
59
+ fast-moving JS plugin API; multi-language headroom via tree-sitter.
60
+
61
+ **Accepted ceiling:** no scope/symbol/import resolution → `const run =
62
+ Effect.runPromise; run(...)` defeats patterns. Rules document this (see D7).
63
+
64
+ ---
65
+
66
+ ## D4 — oxlint/oxc reserved as future *semantic tier*; no abstraction now
67
+
68
+ **Date:** 2026-08-28
69
+
70
+ **Context:** React Doctor's architecture (shared `core`, both eslint-plugin
71
+ and oxlint-plugin adapters) proves the multi-executor endgame. oxc semantic
72
+ analysis (scope, symbols, CFG, type-aware via tsgolint) covers ast-grep's
73
+ ceiling.
74
+
75
+ **Decision:** Do not build an executor abstraction, an oxlint tier, or an SDK
76
+ now. A future rule manifest may declare `executor: ast-grep | oxlint` with a
77
+ shared fixture format — but only when evidence demands it.
78
+
79
+ **Evidence that triggers D4:** the kill test's false-negative log shows alias
80
+ /import-sensitivity causing systematic misses on real conventions.
81
+
82
+ **Explicitly rejected:** hand-rolling a custom engine on oxc internals
83
+ (React Doctor's path — funded team, full-time, months).
84
+
85
+ ---
86
+
87
+ ## D5 — Fixtures ship with every rule; the fixtures ARE the product
88
+
89
+ **Date:** 2026-08-28
90
+
91
+ **Context:** The hard failure mode of LLM-generated rules is not bad code —
92
+ it's *good code for a neighboring spec* ("inside an Effect workflow" is
93
+ ambiguous between two human experts).
94
+
95
+ **Decision:** Every rule must ship with positive + negative fixtures
96
+ executable by `ast-grep test`. The authoring loop is: generate → run against
97
+ real code → model adjudicates candidate findings → disagreements crystallize
98
+ into fixtures → CI runs the pure rule.
99
+
100
+ **Consequences:** Model-level precision at authoring time; determinism and
101
+ zero inference in CI. The fixtures make saved rules maintainable and give the
102
+ repair loop grounding.
103
+
104
+ ---
105
+
106
+ ## D6 — Evidence-first sequencing; Wayfinder deferred
107
+
108
+ **Date:** 2026-08-28
109
+
110
+ **Context:** Toolbox includes Matt Pocock skills (wayfinder, to-spec,
111
+ prototype). Wayfinder is for work too big for one session, wrapped in fog.
112
+
113
+ **Decision:** Run the [kill test](kill-test.md) first (prototype-shaped:
114
+ throwaway code answering a question). Only then `to-spec` into a real spec.
115
+ Wayfinder earns its place later if scope grows (language #2, semantic tier) —
116
+ the fog will exist then.
117
+
118
+ ---
119
+
120
+ ## D7 — Rule honesty: declared analysis level and blind spots
121
+
122
+ **Date:** 2026-08-28
123
+
124
+ **Context:** Trust in generated rules requires knowing what they *can't* see.
125
+
126
+ **Decision:** Every rule carries metadata declaring its analysis level —
127
+ `syntactic` (pattern), `relational` (within-file structure), `multi-file`
128
+ (import/path-aware) — and known blind spots ("syntactic rule; may miss
129
+ aliased imports"). Reported alongside findings.
130
+
131
+ ---
132
+
133
+ ## D8 — CLI generation delegates to the user's installed agent (adapter pattern)
134
+
135
+ **Date:** 2026-08-28
136
+
137
+ **Context:** The desired UX is "an LLM produces rules on the fly" via a CLI.
138
+ Cloudflare Code Mode was raised again as the substrate. The pattern (an LLM
139
+ writes executable analysis) is the product's core; Cloudflare's *product* is
140
+ not required for it.
141
+
142
+ **Decision:** `any-doctor generate "<intent>"` builds a prompt from the
143
+ generation skill and delegates to whatever agent the user already has:
144
+ `claude -p`, `codex exec`, `opencode run`, or a custom command via
145
+ `--agent` / `ANY_DOCTOR_AGENT` (with `{prompt}` placeholder support). No
146
+ first-party LLM integration, no API keys, no server. After the agent returns,
147
+ the CLI **independently verifies** with `sg test` — an agent's rule is not
148
+ trusted until the deterministic harness passes it.
149
+
150
+ **Consequences:** Keyless OSS that rides the agent ecosystem; generation
151
+ quality improves with the user's agent for free. Generation UX varies by
152
+ agent. The first-shot-yield measurement (the product's core metric) requires
153
+ a working agent install at runtime.
154
+
155
+ ---
156
+
157
+ ## D9 — Code Mode pattern is the core; the artifact is a generated program, not YAML; the CLI report is the product (supersedes D3, amends D2/D8)
158
+
159
+ **Date:** 2026-08-28
160
+
161
+ **Context:** Nick corrected course after running react-doctor on a real repo.
162
+ The product vision: `any-doctor "<what to check>"` → the user's LLM produces
163
+ a static-analysis program (via the Code Mode *pattern*: LLM writes TypeScript
164
+ against our typed SDK — **no Cloudflare dependency**) → the CLI runs it and
165
+ renders a React-Doctor-class report (score header, category rollup,
166
+ rule-grouped findings with file:line evidence) → findings feed the LLM to
167
+ fix. The YAML rules I built optimized for auditability and lost the thread:
168
+ they were lint config, not generated analysis programs.
169
+
170
+ **Decision:**
171
+ - The generated artifact is a **`doctor(ctx)` TypeScript program** against a
172
+ typed SDK (`ctx`: files, parse/query, symbols/imports, report builder).
173
+ Code Mode the pattern; local execution; vendor-free.
174
+ - The **CLI report is a first-class deliverable**, matching the React Doctor
175
+ experience (experienced first-hand on a real production repo: 253 files /
176
+ 104ms, score, categories, grouped evidence).
177
+ - BYO-agent generation (D8) unchanged. Deterministic reruns unchanged: the
178
+ saved program re-executes with zero inference.
179
+ - YAML rules are demoted out of the product surface. The five prototype rules
180
+ become reference intents/specs; ast-grep may remain an internal engine
181
+ primitive behind `ctx`, never the user-facing artifact.
182
+ - Trust layer carries over: fixtures gate generated programs; the
183
+ REPAIR-LOG lessons move into the program-generation prompt.
184
+
185
+ **Consequences:** Bigger build (SDK + isolate/subprocess sandbox + report
186
+ renderer). The falsifiable next milestone: one intent → generated program →
187
+ report on a real repo, measured against a known ground truth.
188
+
189
+ ---
190
+
191
+ ## D10 — YAML-era surfaces deleted, not frozen
192
+
193
+ **Date:** 2026-08-28
194
+
195
+ **Context:** The architecture review (2026-08-28) found the repo's front doors
196
+ still sold the superseded YAML direction. The grilling loop for the doctor
197
+ contract offered freeze vs delete vs rewrite; Nick rejected freezing: "we're
198
+ way too early to be locking in decisions like this — we don't need to keep
199
+ the messy stuff."
200
+
201
+ **Decision:** Deleted outright: the `generate`/`init`/`test`/`scan`/`list`
202
+ CLI commands, the YAML generation skill, the prototype rule pack (rules,
203
+ fixtures, snapshots, sgconfig), and the fake-agent dev script. Knowledge
204
+ salvaged to `docs/` (RESULTS.md, REPAIR-LOG.md, intents.md); the seeded
205
+ sample-app moved to `fixtures/sample-app`. D8's agent-adapter code is
206
+ deleted with it; generation returns (rewritten for doctor programs) once
207
+ the program-generation skill exists.
208
+
209
+ **Consequences:** The CLI is exactly two commands — `run` and `verify` —
210
+ matching what the codebase actually is. Nothing in the repo contradicts the
211
+ decision log. Regeneration of anything deleted is cheap: the decisions log
212
+ plus docs/ hold the rationale, and the kill-test learnings live in
213
+ REPAIR-LOG.md.
214
+
215
+ ---
216
+
217
+ ## D11 — Architecture review round 2: the contract becomes enforcement
218
+
219
+ **Date:** 2026-08-28
220
+
221
+ **Context:** Second cold audit found the pivot's promises had gaps: no
222
+ author-facing types (Q3 unimplemented), spoofable sentinel frame, verify
223
+ losing all results on one crash, seed path traversal, false-green sg parse
224
+ failures, and a repo with zero commits and no .gitignore.
225
+
226
+ **Decision:** Adopted the review's top recommendation. Implemented:
227
+ declarations + `types` entry shipped (doctor authors can now reference
228
+ real types); meta validated against the severity enum; sync `doctor()`
229
+ rejected with a clear message; unparseable sg output is a loud error,
230
+ never an empty finding set; last-sentinel-wins frame parsing; doctor
231
+ stdout forwarded to stderr; per-fixture fault isolation (a crashing
232
+ fixture is a named failing result, siblings still report); seed paths
233
+ rejected unless contained in the verify sandbox. Repo: `.gitignore`
234
+ added, real `npm install` replaces the borrowed-tsc/symlink build,
235
+ build script cross-platform, docs shipped in the tarball, `npm test`
236
+ builds first (pretest). CONTEXT.md corrected where it promised more
237
+ than the code does (isolation is by convention today; fixtures match
238
+ (file, line); determinism is per engine version). The initial commit is
239
+ deliberately left to Nick.
240
+
241
+ ---
242
+
243
+ ## D12 — Generation returns: skill for doctor programs, agent adapter restored
244
+
245
+ **Date:** 2026-08-29
246
+
247
+ **Context:** D10 deleted the YAML-era generation flow. Nick clarified the
248
+ deletion intent (the YAML files specifically) and greenlit restoring what
249
+ the doctor-program experience needs.
250
+
251
+ **Decision:** Restored, rewritten for the program contract: the generation
252
+ skill (`skill/any-doctor.skill.md` — contract, workflow, fixture discipline,
253
+ honesty rules, and the failure lessons from REPAIR-LOG), the
254
+ `any-doctor generate "<intent>"` command (agent adapter: claude/codex/
255
+ opencode, custom via `--agent`/`ANY_DOCTOR_AGENT`; the CLI independently
256
+ re-runs `verify` as the post-generation gate), and `dev/fake-agent.sh` for
257
+ plumbing tests without a real agent. The YAML-era fixtures and rule pack
258
+ remain retired — their role is played by `.fixtures.mjs` files.
259
+
260
+ **Consequences:** The full loop works: intent → agent writes doctor + fixtures
261
+ → verify gate → scan. What is still unmeasured: first-shot yield with a real
262
+ agent (requires a working agent install; plumbing proven via fake agent).
263
+ D8's adapter design is hereby re-instantiated in src/cli.ts.
264
+
265
+ ---
266
+
267
+ ## D13 — Doctor discovery & registry UX (planned; spec in docs/features.md)
268
+
269
+ **Date:** 2026-08-29
270
+
271
+ **Context:** Nick's desired experience: `any-doctor verify` / `run` without
272
+ a doctor argument should offer a fuzzy-searchable picker; doctors created
273
+ by generate should be automatically saved in their own non-interfering
274
+ directory. Not previously documented anywhere.
275
+
276
+ **Decision:** Adopted as the next feature set after generation. The build
277
+ spec lives in `docs/features.md` (F1 no-arg fuzzy selection, F2
278
+ repo-local + user-global registry with auto-save and a non-interference
279
+ rule, F3 related batch commands). Implementation follows that file;
280
+ deviations update it in the same commit.
281
+
282
+ ---
283
+
284
+ ## D14 — Any Doctor equips agents; it never deploys them
285
+
286
+ **Date:** 2026-08-29 (amended 2026-09-03)
287
+
288
+ **Context:** The run menu shipped a "Hand off to an agent" item that spawned
289
+ the user's agent, and `generate` spawned the agent headlessly. Nick
290
+ corrected the model: React Doctor's actual pattern is copy-the-findings;
291
+ Any Doctor's only relationship to agents is *equipping* them — the skill
292
+ and the contract — so they can create doctors that fit the interface.
293
+
294
+ **Decision:** No Any Doctor product command ever launches an agent process.
295
+ - The run menu's handoff is now **"Copy findings for your agent"** — a
296
+ ready-to-paste fix prompt on the clipboard (pbcopy/wl-copy/clip).
297
+ - `generate` is **prompt-only**: it plants the skill as `AGENTS.md` in the
298
+ scope dir (agents load it natively), copies the exact generation prompt
299
+ (skill + intent + verify command), and tells the user to run `verify`
300
+ afterward. No spawning, no API keys, no adapter — works with any agent,
301
+ including GUI agents that have no CLI.
302
+ - `dev/first-shot.mjs` remains the sole spawner: it is a measurement tool,
303
+ not product.
304
+ - Deleted: `src/agents.ts`, the agent adapter, and registration-on-generate
305
+ (discovery scans directories; the index is an optional cache).
306
+
307
+ **Consequences:** `generate` is instant and free. Generation quality now
308
+ depends on the skill + the user's agent in the user's own session, which is
309
+ exactly the surface we maintain. The first-shot measurement runs in the
310
+ user's environment by design.
311
+
312
+ ---
313
+
314
+ ## D15 — The npx experience: findings first, creation separate
315
+
316
+ **Date:** 2026-09-05
317
+
318
+ **Context:** React Doctor's traction lesson: `npx react-doctor@latest` gives
319
+ findings in under a minute — no install, no account, no key. Any Doctor's
320
+ cold start inverted that: author doctors first, run later. Nick rethought
321
+ the onboarding with two constraints: no intent routing and no AI-orchestrated
322
+ creation ("weird territory"), and creation must not live inside the running
323
+ experience.
324
+
325
+ **Decision:**
326
+ - `npx any-doctor` with no arguments **runs** instead of printing usage:
327
+ bundled doctors are discovered, the picker appears, findings follow.
328
+ (`help` still prints usage.)
329
+ - **Bundled scope:** a small, curated first-party doctor pack ships inside
330
+ the package. Read-only, lowest priority — repo-local wins collisions,
331
+ then user-global, then bundled. No network at run time.
332
+ - **`init` means adoption, not creation:** copies the bundled pack into
333
+ `./doctors/` (skipping existing files) so a team can edit, prune, and
334
+ commit it. Bundled is a starting point, not a dependency.
335
+ - **`create "<intent>"`** is the creation path — today's prompt-only
336
+ `generate`, renamed to plainer English; `generate` remains as a hidden
337
+ alias. D14's posture is unchanged: no spawning, no keys, the user pastes
338
+ the prompt into their own agent.
339
+ - **Explicitly rejected:** intent routing (matching free-text intents to
340
+ doctors) and any in-CLI agent orchestration or model integration.
341
+
342
+ **Consequences:** The run experience stays pure — select a doctor, see
343
+ findings; nothing pitches AI at a user who hasn't bought in yet. Creation is
344
+ advertised only where curiosity peaks: one dim next-steps line after an
345
+ interactive session, a mention in `init`'s output, and the usage text.
346
+ `run`/`verify` never touch a model (unchanged). Build order: bundled scope →
347
+ no-args-runs → `init` + `create` rename.
348
+
349
+ **Amendment (2026-09-05, post-D16):** Lived experience corrected the flow:
350
+ with the check tree, the dashboard is itself the selection surface, and a
351
+ doctor picker in front of it is friction — the worst doctor lands one enter
352
+ deep instead of on screen. Bare `run` (and the future no-arg npx entry) now
353
+ aggregates every discovered doctor straight into the tree dashboard, no
354
+ picker: doctor rows (worst-severity-first, counts, collapsible) → check
355
+ rows → instances. The picker remains only for `verify`, where choosing one
356
+ doctor to gate is the actual job.
357
+
358
+ ---
359
+
360
+ ## D16 — The doctor experience: category doctors and the check tree
361
+
362
+ **Date:** 2026-09-05
363
+
364
+ **Context:** Nick's product direction: a doctor should own a whole
365
+ category a team cares about ("the async doctor"), not one narrow pattern.
366
+ Opening it shows the *types* of issues as first-class things to select;
367
+ drilling into a type shows where it bites. The contract already allowed
368
+ this (DoctorMeta.checks, finding.rule, the Check glossary term) — the
369
+ pilots just never used it, and the dashboard had no check-level
370
+ navigation.
371
+
372
+ **Decision:**
373
+ - **Flagship:** `doctors/async-doctor.mjs` consolidates fetch-without-
374
+ AbortSignal, unawaited async map, and uncleared setTimeout-in-effect
375
+ as three checks with merged fixtures (15/15 green). The three
376
+ single-pattern originals are deleted.
377
+ - **Dashboard tree:** multi-check doctors render doctor → check →
378
+ instance. Check rows show glyph, description, and instance count;
379
+ enter (or →) expands, ← collapses; enter on an instance copies its
380
+ scoped context as before. Check rows carry a check-level detail pane
381
+ (why/impact/fix/blast radius). Single-check doctors render exactly as
382
+ before — no tree, no behavior change.
383
+ - **Severity-first, after React Doctor's research:** checks sort errors
384
+ before warnings before info (count descending within a band); error
385
+ checks are expanded on entry ("errors always show"); warning/info
386
+ checks start collapsed. An expanded check shows its first 50 instances
387
+ plus a "… and N more — fix a few and re-scan" affordance — the re-scan
388
+ loop is the pagination, not in-session paging.
389
+ - **Keys:** the feed decodes →/← as their own keys alongside ↑/↓.
390
+
391
+ **Consequences:** The picker lists categories, not fragments; one spawn
392
+ covers a whole concern. The dashboard triages: highest severity on
393
+ screen at entry, categories summarized, instances one enter away.
394
+
395
+ ---
396
+
397
+ ## D17 — Confinement: layered runtime enforcement, not a sandbox claim
398
+
399
+ **Date:** 2026-09-05
400
+
401
+ **Context:** Doctors are agent-authored programs executed against a repo.
402
+ The realistic threat is sloppy or prompt-injected code, not a targeted
403
+ attacker; the proportionate answer had to add no OS-level machinery
404
+ (containers, seatbelt profiles — explicitly rejected as over-engineering).
405
+
406
+ **Decision:** Four probe-gated layers, no override in any mode. (1) A
407
+ static capability scan refuses named capabilities before execution —
408
+ tripwire for slop, honest label: it sees the doctor file only. (2) The
409
+ doctor process runs under Node's permission model (writes, subprocesses,
410
+ and native addons denied — worker threads exist only as the import guard's
411
+ carrier and inherit every denial; reads open; verify writes to the temp
412
+ dir only). `--allow-worker` exists solely for the import guard's hook
413
+ thread — denials verified to propagate into worker threads (fs write and
414
+ subprocess are denied inside them). Node prints a SecurityWarning for the
415
+ flag on
416
+ every run; it is suppressed (`--disable-warning=SecurityWarning`) because
417
+ the verified behavior is recorded here instead — the banner would alarm
418
+ every run while changing nothing. (3) An import guard
419
+ (module resolve hook) refuses every module a doctor tries to reach:
420
+ builtins, helper files, npm packages, static, dynamic, or computed — a
421
+ doctor is a single self-contained file (the guard itself stays a
422
+ standalone hand-maintained file in bin/, the one exception to the tsc
423
+ build: a resolve hook may not import anything). (4) The network globals (fetch,
424
+ WebSocket) are deleted from the process; with imports vetoed there is no
425
+ socket API left. `ctx.search` answers over a dedicated channel: the host
426
+ runs ast-grep, the doctor asks. Amends D4: the Engine seam earned a
427
+ module when it had two call sites; the sdk's legacy direct fallback was
428
+ later deleted (a channel-less ctx.search fails loudly instead of running
429
+ ast-grep unconfined), leaving the host as Engine's one caller — the module
430
+ stays for locality: it is where a future backend slots in alone.
431
+
432
+ **Consequences:** Old runtimes degrade to fewer layers (probe-gated, never
433
+ to zero — the static scan is unconditional). Network on future runtimes
434
+ gains a real deny flag eventually; the global strip covers today's.
435
+ Residual honesty: process.env is readable (no outbound channel exists),
436
+ DoS is bounded by the loader timeout, the search root check is lexical (a
437
+ symlink inside the target pointing out is followed — accepted under the
438
+ slop-not-adversary model), and no layered scheme rules out
439
+ engine-internals exotica — the issue #3 tripwire (escalate if third-party
440
+ doctor distribution ships) stays armed.
441
+
442
+ ---
443
+
444
+ ## Open questions
445
+
446
+ - Name: "any-doctor" is a working title.
447
+ - v1 surface: pure skill + harness, or skill + CLI from day one?
448
+ - Timing of the time-aware primitive (git blame / diff-scoped reporting — the
449
+ trick that makes opinionated rules adoptable despite legacy backlogs, same
450
+ as React Doctor's PR-diff CI mode).
@@ -0,0 +1,122 @@
1
+ # Example catalog
2
+
3
+ Brainstormed rule intents showing the breadth of "anything JavaScript
4
+ related." Each is written as the literal sentence you'd type. Tier tags:
5
+
6
+ - `[P]` syntactic — pure pattern match
7
+ - `[R]` relational — within-file structure (`inside`/`has`/`precedes`)
8
+ - `[M]` multi-file — import paths / file graph / repo awareness
9
+ - `[S]` semantic — needs scope/types (future oxlint tier, D4)
10
+
11
+ The rule intents are raw material for the kill test, the demo, and the
12
+ launch README GIF. The pitch is **one tool, one sentence each, unrelated
13
+ frameworks**.
14
+
15
+ ## Async correctness
16
+
17
+ Emptier than people assume — most of this has zero linter coverage.
18
+
19
+ - Find `.map(async ...)` results that are never wrapped in `Promise.all` `[R]`
20
+ - Find `fetch` calls without an `AbortSignal` `[R]` — orphaned requests
21
+ - Find `setTimeout` in effects without a matching `clearTimeout` `[R]`
22
+ - Find `addEventListener` in effects with no `removeEventListener` in
23
+ cleanup `[R]` — paired-symmetry rules
24
+ - Find `await` inside `for` loops `[R]` — exists (`no-await-in-loop`) but
25
+ rarely enabled; the sentence version that suggests `Promise.all` sells
26
+ - ⚠️ Floating promises generally is type-aware territory — typescript-eslint
27
+ owns it; don't lead with it
28
+
29
+ ## AI-era conventions
30
+
31
+ Emptiest field, highest 2026 demand.
32
+
33
+ - Find direct imports of `openai` / `anthropic` outside `lib/ai/` `[P]` —
34
+ gateway enforcement
35
+ - Find LLM responses that get `JSON.parse`'d without schema validation `[R]`
36
+ - Find prompt template literals interpolating raw user input `[R]` —
37
+ prompt-injection hygiene; no linter has this category yet
38
+ - Find `catch` blocks that swallow errors / `catch { return null }` `[P]` —
39
+ framed as "find where the agent hallucinated a swallow"
40
+
41
+ ## Auth / data-layer ordering
42
+
43
+ The screenshot-able demo category.
44
+
45
+ - Find route handlers that touch the DB before calling `requireUser()` `[R]`
46
+ — an *ordering* constraint; `precedes`/`follows` territory
47
+ - Find Prisma queries using the bare `db` client instead of the `tx`
48
+ transaction handle `[R]`
49
+ - Only `api/` may import from `db/` `[M]` — architecture-as-one-sentence
50
+
51
+ ## Temporary migration rules
52
+
53
+ The sleeper hit — write one, run the audit, delete it. Nobody productizes
54
+ these.
55
+
56
+ - Find every remaining import of the old design-system Button `[P]`
57
+ - Find code still reading the deprecated `config.legacyFlags` key `[P]`
58
+ - Find API routes not yet on the new middleware chain `[R]`
59
+ - Find remaining `moment` imports (we migrated to date-fns) `[P]`
60
+
61
+ ## Upgrade doctors (changelog → rules)
62
+
63
+ The batch version of the above: point any-doctor at a library's changelog /
64
+ migration guide **while still on the old version**, and it compiles the
65
+ breaking changes into a set of temporary rules — an upgrade *cost estimate*
66
+ ("340 call sites affected, grouped by breaking change").
67
+
68
+ Key mechanics and caveats (2026-08-28 discussion):
69
+
70
+ - **The compiler overlap:** post-upgrade, `tsc` already catches removed
71
+ exports / renamed APIs / signature mismatches for free in TS repos. Our
72
+ differentiated positions: (a) *pre-upgrade* scanning — the compiler can't
73
+ see v2's rules while you're on v1; (b) semantic changes invisible to
74
+ types; (c) JS repos.
75
+ - **Migration guides are before/after pairs** — the "before" snippets are
76
+ should-flag fixtures, the "after" snippets are should-not-flag fixtures.
77
+ The fixture discipline (D5) is sourced nearly free from the input doc.
78
+ - **Changelog quality is the variable.** Each breaking change gets
79
+ classified: pattern-expressible / relational / needs-types / not-statically-
80
+ checkable — and the report states what was verified vs. what couldn't be
81
+ (D7 honesty). "Error messages changed" → "can't check statically," never
82
+ silently dropped.
83
+
84
+ Worked examples: Express 4→5 (`app.del`, `'*'` → `'*splat'`) `[P]`;
85
+ React 18→19 (`defaultProps` on function components, string refs,
86
+ `propTypes`) `[P]`; Zod 3→4 (`z.record` now requires two args) `[P]`.
87
+
88
+ ## House style (the Dylan Mulroy category)
89
+
90
+ Opinionated rules that could never live in a public plugin — they'd be noise
91
+ for everyone else.
92
+
93
+ - Never import across layers via relative `../..` — use `@/` aliases `[M]`
94
+ - Never construct SQL by string concatenation `[P]`
95
+ - Never pass `Date` objects across the API boundary — ISO strings only `[R]`
96
+ - Server actions must call `requireUser()` before data access `[R]`
97
+
98
+ ## Inventory / audit mode
99
+
100
+ The scanner, not the guard — recall matters, humans triage.
101
+
102
+ - List every external HTTP endpoint we call — gateway-building inventory
103
+ - List every route handler and whether it has auth `[R]`
104
+ - List every place we bypass the logger `[R]`
105
+
106
+ ## Future harness primitive: time-aware rules
107
+
108
+ Add a `git blame` / diff primitive and these become possible:
109
+
110
+ - Flag `as any` and non-null assertions added *after* the AI-pilot rollout
111
+ - Report only violations introduced by this PR (React Doctor's CI trick —
112
+ the adoption unlock for opinionated rules despite legacy backlogs)
113
+
114
+ See [decisions.md](decisions.md) open questions for timing.
115
+
116
+ ## Effect-specific starters (dogfood domain)
117
+
118
+ - Find `Effect.runPromise` calls inside Effect workflows `[S]` — honest tag:
119
+ needs the semantic tier to avoid false positives on the legitimate
120
+ program-boundary call; a syntactic version will document that blind spot
121
+ - Find `Date.now()` inside Effect code — suggest `Clock` `[R]`
122
+ - Find raw promise construction wrapping Effect operations `[R]`
@@ -0,0 +1,67 @@
1
+ # Planned features — doctor discovery & registry
2
+
3
+ Status: F1, F2, and F3 are BUILT on the `buildout` branch — discovery
4
+ scopes, the fuzzy picker (now `verify`-only: bare `run` aggregates every
5
+ doctor straight into the review tree, per D15's amendment), both `--all`
6
+ batch modes, and the repo-local-wins collision policy (with an origin
7
+ suffix when one doctor shadows another). The index.json registration cache
8
+ was dropped (D14): the directory is the registry.
9
+
10
+ Source: Nick's direction (2026-08-29) — "when someone types any-doctor
11
+ verify or any-doctor run without specifying which doctor, there should be
12
+ a select for it, searchable with fuzzy search; and doctors that get
13
+ created should be automatically stored somewhere — its own little doctors
14
+ directory that doesn't interfere."
15
+
16
+ ## F1 — No-argument interactive selection
17
+
18
+ When `run` or `verify` is invoked WITHOUT a doctor path:
19
+
20
+ - **TTY session:** show a fuzzy-searchable picker. Type to filter,
21
+ ↑↓ to move, enter to select, esc/q to cancel. Picker lists every
22
+ discovered doctor by its one-line description and severity glyph; the
23
+ id feeds fuzzy matching and appears in the report and review browser.
24
+ On select: proceed exactly as if the path had been typed
25
+ (`verify` → fixture gate; `run` → report + findings browser).
26
+ - **Non-TTY (piped/CI):** never prompt. Print the discovered doctor list
27
+ and exit non-zero with a clear "specify a doctor" message.
28
+ - Precedence for discovery: `./doctors/` (repo-local) first, then the
29
+ user registry (F2), each entry labeled with its origin.
30
+
31
+ ## F2 — Doctor registry & auto-save
32
+
33
+ Any doctor created by `generate` is automatically **saved** — registered
34
+ so the picker in F1 can find it later. Two scopes, no mixing:
35
+
36
+ - **Repo-local:** `./doctors/` — already the format. A doctor lives next
37
+ to its fixtures; committing both to the consuming repo is encouraged.
38
+ - **User-global:** `~/.any-doctor/doctors/` — for doctors a user wants
39
+ available in every repo. `generate --global` writes here; a doctor
40
+ here is usable from any directory.
41
+
42
+ Registration data lives beside the doctor: its `meta` is the registry
43
+ entry (id, description, severity, blindSpots). The directory is the
44
+ whole registry — there is no index file; discovery scans the directory.
45
+
46
+ **Non-interference rule:** discovery and registry reads never write to
47
+ the target repo being scanned. `run` reads doctors and code; it writes
48
+ nothing anywhere. `verify` writes only to its own temp sandbox.
49
+
50
+ ## F3 — Related (noted, not designed)
51
+
52
+ - `run --all` — execute every discovered doctor, one combined report.
53
+ - `verify --all` — fixture-gate every discovered doctor; exit non-zero
54
+ if any fails.
55
+ - Registry collision policy: same slug in both scopes → repo-local wins,
56
+ picker shows the origin suffix.
57
+ - Sharing/publishing doctor packs — out of scope until F1/F2 exist.
58
+
59
+ ## Implementation notes
60
+
61
+ - Fuzzy matching must be a small zero-dep scorer (subsequence match,
62
+ prefer consecutive runs and word starts — same spirit as React
63
+ Doctor's fuzzy-match.ts, which predates this spec).
64
+ - The picker is raw-mode like the findings browser; frame-builder stays
65
+ pure and golden-testable, per the architecture review.
66
+ - New glossary terms go into CONTEXT.md when these land: "registry",
67
+ "scope" (repo-local vs user-global).
@@ -0,0 +1,18 @@
1
+ # First-shot yield results
2
+
3
+ Run: 2026-08-29T23:16:14.497Z · agent: opencode · gate: any-doctor verify
4
+
5
+ **10/10 passed so far (0 remaining)**
6
+
7
+ | result | intent | detail |
8
+ |---|---|---|
9
+ | PASS | Find fetch calls without an AbortSignal | |
10
+ | PASS | Find .map(async ...) results that are never wrapped in Promise.all | |
11
+ | PASS | Find empty catch blocks that swallow errors | |
12
+ | PASS | Find direct imports of the openai or anthropic SDK outside lib/ai | |
13
+ | PASS | Find Date.now used inside Effect.gen blocks | |
14
+ | PASS | Find z.record called with a single argument | |
15
+ | PASS | Find route handlers that touch the database before calling requireUser | |
16
+ | PASS | Find setTimeout calls inside useEffect without a matching clearTimeout | |
17
+ | PASS | Find JSON.parse calls on LLM API responses without schema validation | |
18
+ | PASS | Find API route files that do not import the auth middleware | |
@@ -0,0 +1,21 @@
1
+ # Intent-only prompts (experiment input)
2
+
3
+ The rules in `rules/` must be derivable from these one-liners alone. This
4
+ file is written FIRST; rules come after. If a rule needs knowledge not in
5
+ its intent, that's a finding for the repair log.
6
+
7
+ ---
8
+
9
+ 1. All LLM calls must go through our gateway in `lib/ai/`. Find direct
10
+ imports of the `openai` or `@anthropic-ai/sdk` packages anywhere outside
11
+ `lib/ai/`.
12
+
13
+ 2. Find `fetch` calls that don't pass an `AbortSignal`.
14
+
15
+ 3. Find `.map(async ...)` calls whose result is not wrapped in `Promise.all`
16
+ — the promises are created but never awaited.
17
+
18
+ 4. Find single-argument `z.record(...)` calls. Zod v4 requires separate key
19
+ and value type arguments (breaking change in the v4 changelog).
20
+
21
+ 5. Find empty `catch` blocks — errors swallowed silently.