@mastra/factory 0.16.0-alpha.9 → 0.16.1-alpha.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (216) hide show
  1. package/README.md +37 -0
  2. package/dist/auth.js +1 -1
  3. package/dist/auth.js.map +1 -1
  4. package/dist/boards/review.d.ts.map +1 -1
  5. package/dist/boards/review.js +19 -8
  6. package/dist/boards/review.js.map +1 -1
  7. package/dist/boards/work.d.ts.map +1 -1
  8. package/dist/boards/work.js +31 -3
  9. package/dist/boards/work.js.map +1 -1
  10. package/dist/capabilities/intake.d.ts +4 -0
  11. package/dist/capabilities/intake.d.ts.map +1 -1
  12. package/dist/capabilities/version-control.d.ts +15 -0
  13. package/dist/capabilities/version-control.d.ts.map +1 -1
  14. package/dist/factory.d.ts +7 -7
  15. package/dist/factory.d.ts.map +1 -1
  16. package/dist/factory.js +100 -40
  17. package/dist/factory.js.map +1 -1
  18. package/dist/integrations/base.d.ts +3 -0
  19. package/dist/integrations/base.d.ts.map +1 -1
  20. package/dist/integrations/github/integration.d.ts.map +1 -1
  21. package/dist/integrations/github/integration.js +21 -0
  22. package/dist/integrations/github/integration.js.map +1 -1
  23. package/dist/integrations/github/routes.js +2 -1
  24. package/dist/integrations/github/routes.js.map +1 -1
  25. package/dist/integrations/github/sandbox.d.ts +41 -42
  26. package/dist/integrations/github/sandbox.d.ts.map +1 -1
  27. package/dist/integrations/github/sandbox.js +416 -227
  28. package/dist/integrations/github/sandbox.js.map +1 -1
  29. package/dist/integrations/github/webhook.d.ts +2 -5
  30. package/dist/integrations/github/webhook.d.ts.map +1 -1
  31. package/dist/integrations/github/webhook.js +8 -48
  32. package/dist/integrations/github/webhook.js.map +1 -1
  33. package/dist/integrations/gitlab/agent-tools.d.ts +14 -0
  34. package/dist/integrations/gitlab/agent-tools.d.ts.map +1 -0
  35. package/dist/integrations/gitlab/agent-tools.js +42 -0
  36. package/dist/integrations/gitlab/agent-tools.js.map +1 -0
  37. package/dist/integrations/gitlab/api.d.ts +222 -0
  38. package/dist/integrations/gitlab/api.d.ts.map +1 -0
  39. package/dist/integrations/gitlab/api.js +274 -0
  40. package/dist/integrations/gitlab/api.js.map +1 -0
  41. package/dist/integrations/gitlab/default-rules.d.ts +125 -0
  42. package/dist/integrations/gitlab/default-rules.d.ts.map +1 -0
  43. package/dist/integrations/gitlab/default-rules.js +172 -0
  44. package/dist/integrations/gitlab/default-rules.js.map +1 -0
  45. package/dist/integrations/gitlab/integration.d.ts +138 -0
  46. package/dist/integrations/gitlab/integration.d.ts.map +1 -0
  47. package/dist/integrations/gitlab/integration.js +705 -0
  48. package/dist/integrations/gitlab/integration.js.map +1 -0
  49. package/dist/integrations/gitlab/issue-reconciler.d.ts +6 -0
  50. package/dist/integrations/gitlab/issue-reconciler.d.ts.map +1 -0
  51. package/dist/integrations/gitlab/issue-reconciler.js +123 -0
  52. package/dist/integrations/gitlab/issue-reconciler.js.map +1 -0
  53. package/dist/integrations/gitlab/merge-request-reconciler.d.ts +6 -0
  54. package/dist/integrations/gitlab/merge-request-reconciler.d.ts.map +1 -0
  55. package/dist/integrations/gitlab/merge-request-reconciler.js +267 -0
  56. package/dist/integrations/gitlab/merge-request-reconciler.js.map +1 -0
  57. package/dist/integrations/gitlab/reconciler.d.ts +5 -0
  58. package/dist/integrations/gitlab/reconciler.d.ts.map +1 -0
  59. package/dist/integrations/gitlab/reconciler.js +31 -0
  60. package/dist/integrations/gitlab/reconciler.js.map +1 -0
  61. package/dist/integrations/gitlab/reconciliation-config.d.ts +3 -0
  62. package/dist/integrations/gitlab/reconciliation-config.d.ts.map +1 -0
  63. package/dist/integrations/gitlab/reconciliation-config.js +17 -0
  64. package/dist/integrations/gitlab/reconciliation-config.js.map +1 -0
  65. package/dist/integrations/gitlab/routes.d.ts +20 -0
  66. package/dist/integrations/gitlab/routes.d.ts.map +1 -0
  67. package/dist/integrations/gitlab/routes.js +473 -0
  68. package/dist/integrations/gitlab/routes.js.map +1 -0
  69. package/dist/integrations/gitlab/rules.d.ts +35 -0
  70. package/dist/integrations/gitlab/rules.d.ts.map +1 -0
  71. package/dist/integrations/gitlab/rules.js +391 -0
  72. package/dist/integrations/gitlab/rules.js.map +1 -0
  73. package/dist/integrations/gitlab/session-subscriptions.d.ts +32 -0
  74. package/dist/integrations/gitlab/session-subscriptions.d.ts.map +1 -0
  75. package/dist/integrations/gitlab/session-subscriptions.js +192 -0
  76. package/dist/integrations/gitlab/session-subscriptions.js.map +1 -0
  77. package/dist/integrations/gitlab/subscriptions.d.ts +70 -0
  78. package/dist/integrations/gitlab/subscriptions.d.ts.map +1 -0
  79. package/dist/integrations/gitlab/subscriptions.js +76 -0
  80. package/dist/integrations/gitlab/subscriptions.js.map +1 -0
  81. package/dist/integrations/gitlab/version-control.d.ts +16 -0
  82. package/dist/integrations/gitlab/version-control.d.ts.map +1 -0
  83. package/dist/integrations/gitlab/version-control.js +622 -0
  84. package/dist/integrations/gitlab/version-control.js.map +1 -0
  85. package/dist/integrations/gitlab/webhook-dispatch.d.ts +75 -0
  86. package/dist/integrations/gitlab/webhook-dispatch.d.ts.map +1 -0
  87. package/dist/integrations/gitlab/webhook-dispatch.js +235 -0
  88. package/dist/integrations/gitlab/webhook-dispatch.js.map +1 -0
  89. package/dist/integrations/gitlab/webhook.d.ts +72 -0
  90. package/dist/integrations/gitlab/webhook.d.ts.map +1 -0
  91. package/dist/integrations/gitlab/webhook.js +201 -0
  92. package/dist/integrations/gitlab/webhook.js.map +1 -0
  93. package/dist/integrations/incidentio/integration.js +1 -1
  94. package/dist/integrations/issue-reconciler.d.ts +3 -3
  95. package/dist/integrations/issue-reconciler.d.ts.map +1 -1
  96. package/dist/integrations/issue-reconciler.js +22 -2
  97. package/dist/integrations/issue-reconciler.js.map +1 -1
  98. package/dist/integrations/platform/github/integration.d.ts.map +1 -1
  99. package/dist/integrations/platform/github/integration.js +21 -0
  100. package/dist/integrations/platform/github/integration.js.map +1 -1
  101. package/dist/integrations/platform/gitlab/event-worker.d.ts +64 -0
  102. package/dist/integrations/platform/gitlab/event-worker.d.ts.map +1 -0
  103. package/dist/integrations/platform/gitlab/event-worker.js +306 -0
  104. package/dist/integrations/platform/gitlab/event-worker.js.map +1 -0
  105. package/dist/integrations/platform/gitlab/integration.d.ts +57 -0
  106. package/dist/integrations/platform/gitlab/integration.d.ts.map +1 -0
  107. package/dist/integrations/platform/gitlab/integration.js +125 -0
  108. package/dist/integrations/platform/gitlab/integration.js.map +1 -0
  109. package/dist/integrations/platform/incidentio/integration.js +1 -1
  110. package/dist/integrations/slack/integration.d.ts.map +1 -1
  111. package/dist/integrations/slack/integration.js +3 -3
  112. package/dist/integrations/slack/integration.js.map +1 -1
  113. package/dist/integrations/slack/slack.d.ts +2 -0
  114. package/dist/integrations/slack/slack.d.ts.map +1 -1
  115. package/dist/integrations/slack/slack.js +21 -7
  116. package/dist/integrations/slack/slack.js.map +1 -1
  117. package/dist/integrations/subscription-session.d.ts +37 -0
  118. package/dist/integrations/subscription-session.d.ts.map +1 -0
  119. package/dist/integrations/subscription-session.js +83 -0
  120. package/dist/integrations/subscription-session.js.map +1 -0
  121. package/dist/packages/_internals/workspace/dist/index.js +1 -11
  122. package/dist/packages/_internals/workspace/dist/index.js.map +1 -1
  123. package/dist/routes/intake.d.ts.map +1 -1
  124. package/dist/routes/intake.js +4 -2
  125. package/dist/routes/intake.js.map +1 -1
  126. package/dist/routes/oauth.js +1 -1
  127. package/dist/routes/projects.d.ts +1 -0
  128. package/dist/routes/projects.d.ts.map +1 -1
  129. package/dist/routes/projects.js +1 -0
  130. package/dist/routes/projects.js.map +1 -1
  131. package/dist/routes/source-control-sessions.d.ts +20 -0
  132. package/dist/routes/source-control-sessions.d.ts.map +1 -0
  133. package/dist/routes/source-control-sessions.js +319 -0
  134. package/dist/routes/source-control-sessions.js.map +1 -0
  135. package/dist/routes/source-control-settings.d.ts +10 -0
  136. package/dist/routes/source-control-settings.d.ts.map +1 -0
  137. package/dist/routes/source-control-settings.js +87 -0
  138. package/dist/routes/source-control-settings.js.map +1 -0
  139. package/dist/routes/surface.d.ts +3 -3
  140. package/dist/routes/surface.d.ts.map +1 -1
  141. package/dist/routes/surface.js +55 -13
  142. package/dist/routes/surface.js.map +1 -1
  143. package/dist/routes/tenant-credentials.d.ts +8 -0
  144. package/dist/routes/tenant-credentials.d.ts.map +1 -1
  145. package/dist/routes/tenant-credentials.js +20 -3
  146. package/dist/routes/tenant-credentials.js.map +1 -1
  147. package/dist/routes/work-items.d.ts.map +1 -1
  148. package/dist/routes/work-items.js +1 -1
  149. package/dist/routes/work-items.js.map +1 -1
  150. package/dist/rules/dispatcher.d.ts.map +1 -1
  151. package/dist/rules/dispatcher.js +7 -1
  152. package/dist/rules/dispatcher.js.map +1 -1
  153. package/dist/rules/start-coordinator.d.ts +1 -1
  154. package/dist/rules/start-coordinator.d.ts.map +1 -1
  155. package/dist/rules/start-coordinator.js +3 -2
  156. package/dist/rules/start-coordinator.js.map +1 -1
  157. package/dist/rules/transition-service.d.ts.map +1 -1
  158. package/dist/rules/transition-service.js +6 -3
  159. package/dist/rules/transition-service.js.map +1 -1
  160. package/dist/rules/types.d.ts +67 -3
  161. package/dist/rules/types.d.ts.map +1 -1
  162. package/dist/rules/types.js +23 -2
  163. package/dist/rules/types.js.map +1 -1
  164. package/dist/rules/validation.d.ts.map +1 -1
  165. package/dist/rules/validation.js +2 -0
  166. package/dist/rules/validation.js.map +1 -1
  167. package/dist/sandbox/git-ref.d.ts +7 -0
  168. package/dist/sandbox/git-ref.d.ts.map +1 -0
  169. package/dist/sandbox/git-ref.js +15 -0
  170. package/dist/sandbox/git-ref.js.map +1 -0
  171. package/dist/sandbox/session-retirement.js +1 -1
  172. package/dist/sandbox/workdir.d.ts.map +1 -1
  173. package/dist/sandbox/workdir.js +6 -6
  174. package/dist/sandbox/workdir.js.map +1 -1
  175. package/dist/session/factory-session.d.ts +31 -1
  176. package/dist/session/factory-session.d.ts.map +1 -1
  177. package/dist/session/factory-session.js +60 -1
  178. package/dist/session/factory-session.js.map +1 -1
  179. package/dist/session/source-control-tools.d.ts +16 -0
  180. package/dist/session/source-control-tools.d.ts.map +1 -0
  181. package/dist/session/source-control-tools.js +472 -0
  182. package/dist/session/source-control-tools.js.map +1 -0
  183. package/dist/storage/domains/intake/base.d.ts +18 -0
  184. package/dist/storage/domains/intake/base.d.ts.map +1 -1
  185. package/dist/storage/domains/intake/base.js +77 -0
  186. package/dist/storage/domains/intake/base.js.map +1 -1
  187. package/dist/work-item-branch.d.ts +3 -1
  188. package/dist/work-item-branch.d.ts.map +1 -1
  189. package/dist/work-item-branch.js +22 -4
  190. package/dist/work-item-branch.js.map +1 -1
  191. package/dist/workspace.d.ts +15 -1
  192. package/dist/workspace.d.ts.map +1 -1
  193. package/dist/workspace.js +57 -23
  194. package/dist/workspace.js.map +1 -1
  195. package/factory-skills/factory-gitlab-rereview/SKILL.md +19 -0
  196. package/factory-skills/factory-gitlab-review/SKILL.md +30 -0
  197. package/factory-skills/factory-rereview/SKILL.md +13 -6
  198. package/factory-skills/factory-review/SKILL.md +13 -6
  199. package/factory-skills/factory-review/references/archaeology.md +225 -0
  200. package/factory-skills/factory-review/references/categories/README.md +34 -0
  201. package/factory-skills/factory-review/references/categories/behavior-change.md +36 -0
  202. package/factory-skills/factory-review/references/categories/bug-fix.md +37 -0
  203. package/factory-skills/factory-review/references/categories/docs.md +32 -0
  204. package/factory-skills/factory-review/references/categories/experimental-flagged.md +34 -0
  205. package/factory-skills/factory-review/references/categories/infra-tooling.md +37 -0
  206. package/factory-skills/factory-review/references/categories/internal-capability.md +35 -0
  207. package/factory-skills/factory-review/references/categories/large-cross-cutting.md +35 -0
  208. package/factory-skills/factory-review/references/categories/mechanical.md +36 -0
  209. package/factory-skills/factory-review/references/categories/performance.md +33 -0
  210. package/factory-skills/factory-review/references/categories/public-api.md +38 -0
  211. package/factory-skills/factory-review/references/categories/refactor.md +35 -0
  212. package/factory-skills/factory-review/references/categories/revert.md +31 -0
  213. package/factory-skills/factory-review/references/categories/schema-storage.md +39 -0
  214. package/factory-skills/factory-review/references/categories/security.md +35 -0
  215. package/factory-skills/factory-review/references/categories/tests-only.md +38 -0
  216. package/package.json +9 -9
@@ -0,0 +1,225 @@
1
+ # Archaeology — command recipes
2
+
3
+ Zero judgment in this file. These are known-working command shapes; adapt them to the repository and preserve only decisive evidence in the review record.
4
+
5
+ **Shell notes.** `gh` output can carry ANSI codes that break `jq`; use `gh`'s built-in `--jq` or prefix with `NO_COLOR=1`. `timeout` is not a command on macOS. Quote paths with spaces.
6
+
7
+ ## Resolve the PR before prediction
8
+
9
+ ```bash
10
+ # Accepts a number, URL, or branch. Keep the first query minimal.
11
+ gh pr view <pr> --json number,title,body,author,baseRefName,url,closingIssuesReferences \
12
+ --jq '{number,title,author:.author.login,base:.baseRefName,url,issues:[.closingIssuesReferences[]?.number],body}'
13
+ ```
14
+
15
+ `baseRefName` is the comparison branch; it may not be `main`. Extract the problem statement from `body` and ignore implementation/change-list sections until after the pre-diff model is written.
16
+
17
+ Owner/repo for later calls:
18
+
19
+ ```bash
20
+ gh repo view --json nameWithOwner --jq .nameWithOwner
21
+ ```
22
+
23
+ ## Open the PR after the pre-diff model
24
+
25
+ ```bash
26
+ # Resolve the head now that isolation is over
27
+ gh pr view <pr> --json headRefName,headRefOid --jq '{head:.headRefName,headSha:.headRefOid}'
28
+
29
+ # Changed files
30
+ gh pr view <pr> --json files --jq '.files[] | "\(.path)\t+\(.additions) -\(.deletions)"'
31
+
32
+ # Commit list
33
+ gh pr view <pr> --json commits --jq '.commits[] | "\(.oid[0:8]) \(.messageHeadline)"'
34
+
35
+ # Reviews and review-thread comments (with resolved state)
36
+ gh api graphql -f query='
37
+ query($owner:String!,$repo:String!,$pr:Int!){
38
+ repository(owner:$owner,name:$repo){ pullRequest(number:$pr){
39
+ reviews(first:50){ pageInfo{hasNextPage endCursor} nodes{ author{login} state submittedAt body } }
40
+ reviewThreads(first:100){ pageInfo{hasNextPage endCursor} nodes{ isResolved isOutdated path line
41
+ comments(first:20){ pageInfo{hasNextPage endCursor} nodes{ author{login} createdAt body } } } }
42
+ }}}' -F owner=<owner> -F repo=<repo> -F pr=<pr>
43
+ # Follow pageInfo.hasNextPage/endCursor on reviews, reviewThreads, and comments until
44
+ # exhausted — a truncated history reads as complete and re-raises omitted findings.
45
+
46
+ # PR-level (issue) comments
47
+ gh api repos/<owner>/<repo>/issues/<pr>/comments --paginate --jq '.[] | "\(.user.login) \(.created_at)\n\(.body)\n---"'
48
+
49
+ # CI status
50
+ gh pr checks <pr> 2>&1 || true
51
+ ```
52
+
53
+ If GraphQL is rate-limited, the REST equivalents (all with `--paginate`): `gh api repos/<owner>/<repo>/pulls/<pr>/files`, `.../pulls/<pr>/commits`, `.../pulls/<pr>/reviews`, `.../pulls/<pr>/comments` (review comments; no resolved state via REST — note that). Check quota with `gh api rate_limit --jq '.resources | {core:.core.remaining, graphql:.graphql.remaining}'`.
54
+
55
+ ## Linked issues (step 2)
56
+
57
+ ```bash
58
+ gh issue view <n> --json number,title,body,author,state,labels,comments \
59
+ --jq '{number,title,author:.author.login,state,labels:[.labels[].name],body,comments:[.comments[] | {author:.author.login,body}]}'
60
+ ```
61
+
62
+ Also scan the PR body for `#NNN`, `fixes`, `closes`, `resolves` — `closingIssuesReferences` only catches the formally linked ones.
63
+
64
+ ## Base-branch worktree before prediction
65
+
66
+ ```bash
67
+ git fetch origin <base>
68
+ git worktree add /tmp/review-base-<pr> origin/<base>
69
+ ```
70
+
71
+ A clean existing base checkout is also fine. Do not open the PR head until the pre-diff model is written.
72
+
73
+ ## Head and diff after prediction
74
+
75
+ ```bash
76
+ # Fork heads are not on origin/<head>; fetch the PR ref and pin the worktree to the head SHA
77
+ gh pr view <pr> --json headRefOid --jq .headRefOid # <sha>
78
+ git fetch origin refs/pull/<pr>/head
79
+ git worktree add --detach /tmp/review-head-<pr> <sha>
80
+ git -C /tmp/review-head-<pr> diff origin/<base>...HEAD --stat
81
+ git -C /tmp/review-head-<pr> diff origin/<base>...HEAD
82
+ ```
83
+
84
+ ## History — when relevant
85
+
86
+ ```bash
87
+ # Recent history of each core file (last 20 commits, not the whole log)
88
+ git log --oneline -20 origin/<base> -- <file>
89
+
90
+ # Who wrote the changed lines, in what commit (run on the base worktree against the pre-PR lines)
91
+ git blame -L <start>,<end> origin/<base> -- <file>
92
+
93
+ # From a blame SHA to its PR
94
+ gh pr list --search "<sha>" --state merged --json number,title,url --jq '.[]'
95
+ # or
96
+ gh api "repos/<owner>/<repo>/commits/<sha>/pulls" --jq '.[] | {number,title,html_url}'
97
+ ```
98
+
99
+ Read the originating PR's description and review thread. That's where the reason lives.
100
+
101
+ ## History — deep (on trigger)
102
+
103
+ ```bash
104
+ # When was a string introduced or removed
105
+ git log -S "<exact string>" --oneline origin/<base> -- <path-or-dot>
106
+
107
+ # Reverts touching the area
108
+ git log --oneline origin/<base> --grep="^Revert" -- <path>
109
+
110
+ # The reverted attempt itself
111
+ gh pr view <n> --json title,body,reviews,comments
112
+
113
+ # Closed-unmerged PRs with the same idea (strongest precedent signal)
114
+ gh pr list --search "<feature terms> is:closed is:unmerged" --state closed --limit 20 --json number,title,closedAt,url --jq '.[]'
115
+
116
+ # Issues describing the same symptom under other words
117
+ gh issue list --search "<symptom terms>" --state all --limit 20 --json number,title,state,url --jq '.[]'
118
+ ```
119
+
120
+ ## Related PRs and issues (when overlap or precedent could matter)
121
+
122
+ ```bash
123
+ # Open PRs touching the same files (merge-order risk)
124
+ gh pr list --state open --limit 50 --json number,title,files --jq '.[] | select(.files[].path | test("<path-fragment>")) | {number,title}'
125
+
126
+ # By feature/function name, open and closed
127
+ gh pr list --search "<term>" --state all --limit 20 --json number,title,state,url --jq '.[]'
128
+ gh issue list --search "<term>" --state all --limit 20 --json number,title,state,url --jq '.[]'
129
+ ```
130
+
131
+ Search terms: the feature name, the file path, the function names touched, the error message from the issue.
132
+
133
+ ## Callers
134
+
135
+ ```bash
136
+ # Every reference to a changed symbol, on the PR branch
137
+ git grep -nE "\b<symbol>\b" -- '*.ts' '*.tsx' '*.js'
138
+
139
+ # Including string references (config keys, docs, dynamic imports)
140
+ git grep -n "<symbol>" -- ':!*.lock'
141
+
142
+ # Near-twin helper search: grep for the *verb* and the *shape*, not the name
143
+ git grep -nE "function (with|create|make)?[A-Za-z]*(Retry|retry)" -- 'packages/*.ts'
144
+ ```
145
+
146
+ Classify every hit: callers to account for versus definitions, comments, documentation, and string references. Read each match before treating it as a caller — a search that also hits the definition or a doc is not evidence of a caller.
147
+
148
+ ## Test on base — when needed
149
+
150
+ Use this only when code reasoning cannot establish whether the regression evidence would fail without the fix and the answer would materially affect confidence. Never replace production files in the current checkout with base versions. Use a disposable worktree; if it cannot be prepared cleanly, stop rather than falling back to modifying the live checkout.
151
+
152
+ ```bash
153
+ ROOT=$(git rev-parse --show-toplevel)
154
+ HEAD=$(gh pr view <pr> --json headRefOid --jq .headRefOid)
155
+ git fetch origin refs/pull/<pr>/head # fork heads are not on origin/<head>
156
+ MB=$(git merge-base origin/<base> "$HEAD")
157
+ WT="${TMPDIR:-/tmp}/review-tob-<pr>-$$"
158
+ test ! -e "$WT" || exit 1
159
+ git -C "$ROOT" worktree add "$WT" "$MB"
160
+ cd "$WT"
161
+
162
+ # Only the test files the PR added or modified (a deleted path does not exist at $HEAD)
163
+ git diff --diff-filter=AM "$MB"..."$HEAD" --name-only -- '**/*.test.*' '**/*.spec.*' '**/__tests__/**' > .test-files
164
+ test -s .test-files || { echo "no test files changed — nothing to run"; }
165
+ xargs -a .test-files git checkout "$HEAD" --
166
+
167
+ # Build what the tests need, using the repo's documented shape (AGENTS.md / CONTRIBUTING). Do not improvise build commands.
168
+ # Inspect package.json scripts and lockfiles first; run everything with GH_TOKEN and GITHUB_TOKEN unset (the PR's code runs here).
169
+ # Then run only those files, keeping the test command's own exit status (pipefail so tail cannot mask a failure):
170
+ set -o pipefail
171
+ test -s .test-files && env -u GH_TOKEN -u GITHUB_TOKEN xargs -a .test-files pnpm vitest run --reporter=dot --bail 1 2>&1 | tail -40
172
+
173
+ cd "$ROOT"
174
+ git worktree remove --force "$WT"
175
+ ```
176
+
177
+ Expected: red. Green means the test doesn't catch the regression — that's a finding, not a success. Record the decisive assertion, not the full output.
178
+
179
+ ## Run the thing (public-api, behavior-change, docs)
180
+
181
+ A scratch project that consumes the package through its public export, not source imports. Build the package per the repo's documented command first (send build output to a separate log, not the transcript). This executes the PR's code: inspect its scripts and lockfile first and run with GH_TOKEN and GITHUB_TOKEN unset.
182
+
183
+ ```bash
184
+ mkdir -p /tmp/review-run-<pr> && cd /tmp/review-run-<pr>
185
+ # link or install the workspace package, write a 10-line script that exercises the claim, run it
186
+ env -u GH_TOKEN -u GITHUB_TOKEN node demo.mjs 2>&1 | tee /tmp/review-run-<pr>/with.txt
187
+ ```
188
+
189
+ For behavior changes, run the same script against a base build into `without.txt`. Summarize the observed difference in the review record; retain transcripts only when they help reproduce it.
190
+
191
+ ## Changeset
192
+
193
+ ```bash
194
+ git diff origin/<base>...HEAD --name-only -- .changeset/
195
+ cat .changeset/*.md # on the branch; read the level and the packages
196
+ ```
197
+
198
+ ## CODEOWNERS / who likely knows
199
+
200
+ ```bash
201
+ # Owners for the touched paths
202
+ grep -n "<path-fragment>" .github/CODEOWNERS 2>/dev/null
203
+ # Or the people who've touched the file most recently
204
+ git shortlog -sn --since="1 year ago" origin/<base> -- <file> | head -5
205
+ ```
206
+
207
+ Name a person under **Open questions**, not "someone."
208
+
209
+ ## Long memory — recovery searches
210
+
211
+ | Memory | What the reviewer "just knows" | Recover from |
212
+ | ----------- | --------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
213
+ | Decisions | "We chose X over Y because…" | Originating PRs (blame → PR), nearest AGENTS.md, design docs, `git grep -nE "because \| workaround \| don't \| see #" -- <files>` near the changed lines |
214
+ | Scars | "Last time someone touched this it broke Z" | `git log --grep="^Revert" -- <path>`; test names that encode bugs (`git grep -nE "should not \| regression \| double \| race" -- <test files>`); `gh issue list --search "<area> regression"` |
215
+ | Users | "Half the community calls this with a string" | `git grep -n "<api>" -- examples/ docs/`; `gh issue list --search "<api>"`; call sites in the repo |
216
+ | Conventions | "We always go through the storage abstraction here" | The sibling feature; the three closest public neighbors; the package `CHANGELOG.md` for recent movement in the area |
217
+
218
+ ## Cleanup
219
+
220
+ ```bash
221
+ git worktree remove --force /tmp/review-base-<pr> 2>/dev/null
222
+ git worktree remove --force /tmp/review-head-<pr> 2>/dev/null
223
+ ```
224
+
225
+ Leave the handoff artifact in place so a re-review can reuse the recorded design and earlier findings.
@@ -0,0 +1,34 @@
1
+ # Category pages
2
+
3
+ One page per common PR category. After recording your own design and before opening the diff, load the pages relevant to the problem. As the actual change reveals additional categories, load those pages before reviewing their portions in depth. Do not skip relevant pages because the main skill seems sufficient. Use their review focus, traps, and approval criteria; reading is required, while checks are selected according to the change and its risks. Every page has the same shape:
4
+
5
+ - **Reviewing** — the _object_ under review. It's not the diff; it's a causal claim, an interface, an equivalence proof, a set of numbers. Same diff format, different thing under scrutiny.
6
+ - **Read first** — what to read before the diff. Reading the diff first is the wrong first move for almost every category.
7
+ - **Done means** — what has to be true for you to approve. This is the bar the category contributes.
8
+ - **Trap** — the characteristic way this category fools a reviewer. Knowing the trap is most of the page's value.
9
+ - **Attention** — where scrutiny goes. Some categories are 80% interface / 20% implementation; some the inverse.
10
+ - **Questions** — the activated subset of the question bank, written out.
11
+ - **Signals → branches** — when you see X, do Y. Opens only when the signal fires.
12
+ - **Verify** — useful ways to falsify the important claims. Recipes in `../archaeology.md`.
13
+
14
+ Consider every entry below against the behavior, compatibility boundaries, and failure modes the change touches, not just the PR's headline category. These pages contain broadly useful review knowledge; their names are not exclusive labels, and the triggers below are examples, not exclusions. If a page could plausibly help, err on the side of reading it. Reading more guidance does not require running every check it describes.
15
+
16
+ PRs are often mixed. Keep materially different conclusions separate. "Should be split" is a finding when mixing makes a portion harder to review than it would be alone; do not average distinct conclusions into one.
17
+
18
+ | Page | Reviewing | When relevant |
19
+ | ------------------------- | ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
20
+ | `mechanical.md` | That it's _only_ that | Renames, formatting, generated files, lockfiles, or version bumps whose noise could hide substantive changes. |
21
+ | `bug-fix.md` | A causal claim | Correctness fixes, regressions, defensive patches, or claims that an observed symptom is resolved. |
22
+ | `behavior-change.md` | Impact on everyone who relied on the old behavior | Changed defaults, ordering, thresholds, errors, retries, or other observable behavior—even with no API signature change. |
23
+ | `internal-capability.md` | Architectural fit | New internal mechanisms or integration with existing layers, hooks, registries, or lifecycle handling. |
24
+ | `public-api.md` | The interface | New or changed public signatures, options, return types, errors, exports, or cross-package APIs. |
25
+ | `refactor.md` | An equivalence proof | Moved or restructured code, replaced implementations, or cleanup that claims to preserve behavior—even inside a feature or fix. |
26
+ | `schema-storage.md` | Irreversibility and the upgrade path | Persisted data, migrations, wire/config formats, or compatibility between independently upgraded packages—even without a schema migration. |
27
+ | `performance.md` | Numbers | Performance claims, hot-path changes, caches, allocations, or work that grows with input size. |
28
+ | `infra-tooling.md` | Blast radius on every developer | Build/CI/scripts, install or release behavior, dependencies, published files, or package exports. |
29
+ | `security.md` | Completeness | Untrusted input, permissions, secrets, filesystem/network trust boundaries, or vulnerability fixes—not just security-labeled PRs. |
30
+ | `docs.md` | Truth | Changed docs or examples, or API/behavior changes that could leave existing documentation false. |
31
+ | `tests-only.md` | Whether they'd catch anything | New, changed, weakened, or removed tests, mocks, fixtures, or shared test helpers—even alongside production changes. |
32
+ | `revert.md` | That it's clean and the reason is recorded | Full or partial reversions, manual undoing of earlier work, or rollback changes with dependent work still present. |
33
+ | `large-cross-cutting.md` | Whether it should be one PR | Multiple concerns, packages, or architectural boundaries; mixed changes whose interactions or structure make review difficult. |
34
+ | `experimental-flagged.md` | Isolation | Experiments, feature flags, opt-in execution paths, or claims that disabling a feature leaves existing behavior untouched. |
@@ -0,0 +1,36 @@
1
+ # Behavior change (no new API)
2
+
3
+ A default, an ordering, a threshold, a format, a timing — something existing users observe changes, without a new surface.
4
+
5
+ **Reviewing:** impact on everyone who relied on the old behavior.
6
+ **Read first:** the callers and consumers of the changed thing — on the base, before the diff.
7
+ **Done means:** every affected caller is accounted for; the changeset level matches (a changed default is breaking); a changelog line exists that a user would understand.
8
+ **Trap:** labeled `patch`; nobody grepped the callers.
9
+ **Attention:** the consumers, not the change. The change is usually small; the blast radius isn't.
10
+
11
+ ## Questions
12
+
13
+ - **Who relied on the old behavior?** Grep every consumer — in the repo, in `examples/`, in docs, in issues that mention the API. The author fixed it for their case.
14
+ - **Was the old behavior a decision or an accident?** `blame` the lines, open the originating PR, read why. If it was a decision, the PR must argue with it, not replace it. If the author doesn't know it was a decision, they haven't.
15
+ - **What happens to existing users on upgrade?** Walk it: a user on the previous version with existing config/code/data updates. What silently changes? Silent behavior changes are worse than errors — nothing tells the user.
16
+ - **Is the changeset level honest?** A `patch` that changes a default is a breaking change wearing a disguise.
17
+ - **What would the changelog line say?** If a user reading it wouldn't know whether they're affected, it's not written yet.
18
+ - **Does the description match the code?** Behavior changes are the most common thing smuggled into PRs labeled "fix" or "refactor."
19
+ - **Is there a compatibility path?** A flag, a deprecation window, an opt-in — or is it a hard cut? If hard, is that justified?
20
+ - **Cross-package.** If the changed behavior is observed across a package boundary, what does the other side see when only one of them upgrades?
21
+ - **In-flight work.** Other open PRs that depend on the old behavior.
22
+
23
+ ## Signals → branches
24
+
25
+ - Category was triaged as "bug fix" but the diff changes a default/order/threshold → re-categorize; the description undersells scope — finding
26
+ - Changeset says `patch` → check whether any consumer's observable output changes; if yes, request the correct level before merge
27
+ - Originating PR explains the old behavior → deep history; the PR must engage with that reasoning
28
+ - No caller grep in the description → do it; every hit is a question
29
+
30
+ ## Verify
31
+
32
+ - **Run it, before and after.** Write a minimal script that exercises the changed behavior through the public surface. Run on base (old behavior) and on the branch (new), then record the decisive difference.
33
+ - Grep callers/consumers on the branch: `archaeology.md` → Callers.
34
+ - `git blame -L` the changed lines on the base; open the originating PR.
35
+ - `ls .changeset/` and read the level.
36
+ - `gh pr list --search` for open PRs touching the same files.
@@ -0,0 +1,37 @@
1
+ # Bug fix
2
+
3
+ **Reviewing:** a causal claim — "X was the cause, and this removes it."
4
+ **Read first:** the issue and its repro → the test → the diff. Never the diff first.
5
+ **Done means:** you believe the cause; the fix lives at the cause; the evidence distinguishes the regression from working behavior.
6
+ **Trap:** symptom patch; "why now" unanswered.
7
+ **Attention:** the cause, not the code. The code is judged by whether it addresses the cause you independently arrived at.
8
+
9
+ ## Questions
10
+
11
+ - **Where does the value first become wrong?** The tell is _where_ the fix lives relative to where the bug was _observed_. A fix at the observation site — a null check where it crashed, a `?.`, a re-sort before display, a retry around the failing call — is almost always a symptom patch. Ask _why_ was the value null / out of order / failing, and follow the data backwards until you find where it _became_ wrong. That's where the fix belongs.
12
+ - **Does this make the bug impossible, or just unobserved here?** If the same bad state can reach another consumer, it's a patch.
13
+ - **Why now?** Bugs have a cause. Either the code was always wrong (then why did it work before — what assumption changed?) or something changed (which commit — `git log -S`, bisect). A real fix can name the cause. "It was broken so I made it not broken" can't.
14
+ - **Does your theory match theirs?** Your prediction included your own theory of the cause. If the author's fix is somewhere other than where your theory says the value goes wrong, one of you is wrong, and finding out which is the review's first job.
15
+ - **Is the fix more defensive, or more correct?** More checks, more fallbacks, more catch blocks — correct code has fewer branches, not more. A fix that adds branches is a symptom patch until proven otherwise.
16
+ - **Did the fix turn a loud failure into a silent one?** A swallowing try/catch or a masking default that makes the symptom disappear without fixing the state.
17
+ - **Would the test have failed before the fix?** For a bug fix this is the central test question. Establish the counterfactual from the assertion and base implementation; code reasoning is enough when the answer is deterministic. Green on base means the test proves nothing, but executing on base is useful only when the answer remains uncertain.
18
+ - **Does the test name match the claim?** "Fixes ordering under concurrent writes" → is there a test with concurrent writes, or one write that checks order? The gap between claim and test scope is where the bugs live.
19
+ - **What does the assertion actually assert?** `toBeDefined()`, `toHaveBeenCalled()`, `not.toThrow()` pass for almost any implementation. A real test asserts the specific value or behavior.
20
+ - **Is the code under test the real code?** If the thing being fixed is mocked out, the test is testing the mock.
21
+ - **Who else calls the changed function?** The author fixed it for their caller. Do the others still work?
22
+
23
+ ## Signals → branches
24
+
25
+ - Fix lives at the observation site (null check, `?.`, retry, re-sort) → trace the data backwards to where it becomes wrong; that's the finding
26
+ - Fix adds branches rather than removing them → symptom patch until proven otherwise
27
+ - Description can't name the cause, only the symptom → "why now" is unanswered; go deep on history
28
+ - Old, stable code → deep history; something changed, find the commit
29
+ - Test asserts weakly or mocks the fixed module → run test-on-base; expect it to be green (i.e. useless)
30
+ - The same pattern (e.g. the same null check) appears in more than one place in the diff → the fix is at the consumers, not the source
31
+
32
+ ## Verify
33
+
34
+ - **test-on-base** (`archaeology.md` → Test on base): use it when the regression test is central and code reasoning cannot establish whether it distinguishes the bug—for example, weak assertions, mocked paths, concurrency, integration behavior, or multiple code paths that could satisfy it. Green means it does not distinguish the regression; whether that blocks depends on what other evidence proves the fix.
35
+ - Use `git log -S "<removed or changed string>"` when a deletion or condition change has unclear history.
36
+ - Account for callers whose behavior could change.
37
+ - If the issue has a useful repro, run it on base and on the branch and record the observed difference.
@@ -0,0 +1,32 @@
1
+ # Docs
2
+
3
+ **Reviewing:** truth. Not prose quality — whether the claims match the code.
4
+ **Read first:** the code the docs describe, on the base (or on the branch if the docs accompany a code change).
5
+ **Done means:** material claims match the code; new or changed examples run; nothing that was true is now stated falsely.
6
+ **Trap:** reviewing prose quality instead of correctness.
7
+ **Attention:** examples and API claims. Adjectives don't matter; signatures, defaults, and behavior do.
8
+
9
+ ## Questions
10
+
11
+ - **Does each claim match the code?** Every "defaults to X," "returns Y," "throws when Z" — find the line that makes it true. Docs drift from code silently and the drift is what users hit.
12
+ - **Do the examples run?** Copy each example into a scratch file, run it against the package. An example that doesn't compile is worse than no example.
13
+ - **Is anything now stale?** If the docs accompany a code change, what _else_ in the docs referenced the old behavior? Grep the docs tree for the old name/default.
14
+ - **Does it describe what, or why?** "Set `foo` to enable bar" is what. Users need "set `foo` when X; leave it off when Y because Z." Missing _why_ isn't a correctness failure, but it's the difference between docs that get read and docs that get skipped.
15
+ - **Would a user find it?** Is it in the place a user would look — next to its siblings, linked from the index?
16
+ - **Is it complete for the surface?** A new option documented without its interaction with existing options; an error case not mentioned.
17
+ - **Terminology.** Same terms as the rest of the docs and the code? A doc that calls it a "session" when the code says "thread" creates a support ticket.
18
+ - **Prose only if it's wrong.** Typos and phrasing are nits at most. Don't spend the review on them.
19
+
20
+ ## Signals → branches
21
+
22
+ - Example with an import → run it; imports are where examples rot first
23
+ - A default or return type stated → find it in the code; if you can't in 30 seconds, it's probably wrong
24
+ - Docs accompany a rename or default change → grep the whole docs tree for the old term
25
+ - New feature docs with no mention of failure modes → check whether the feature has any; if it does, request that they be documented
26
+
27
+ ## Verify
28
+
29
+ - Run new or changed examples in a scratch project through the public package surface; sample unchanged examples only when broad edits create uncertainty.
30
+ - For each material default/return/throw claim, find the supporting code.
31
+ - `grep -rn "<old term>" docs/` on the branch if anything was renamed.
32
+ - If the docs site has a build (`pnpm build` in the docs package), run it — broken links and MDX errors surface there.
@@ -0,0 +1,34 @@
1
+ # Experimental / behind a flag
2
+
3
+ **Reviewing:** isolation. The bar on polish is lower; the bar on containment is higher.
4
+ **Read first:** the flag boundary — where the flag is checked, and everything on the "off" side of it.
5
+ **Done means:** off by default; zero effect when off; removable without surgery; a plan for graduating or deleting it.
6
+ **Trap:** the flag leaks — "off" still changes something; "temporary" code that stays forever.
7
+ **Attention:** the off path. The on path is experimental by declaration; the off path is production.
8
+
9
+ ## Questions
10
+
11
+ - **Is it actually off by default?** Find the default. Find every place the flag is read. Config, env var, constructor option — one of them defaulting to on is a leak.
12
+ - **Zero effect when off?** Trace the off path. New imports still load; new constructor args still validate; new fields still serialize. Any of these is an effect when off.
13
+ - **Where is the flag checked?** One place (good) or scattered through the codebase (each check is a place the flag can be forgotten)?
14
+ - **Is it removable?** When the experiment ends, can the flag and the off-path be deleted in one PR without touching unrelated code? If the flag check is woven into existing logic, it isn't.
15
+ - **Is there a plan?** Graduation criteria, an owner, a date, or an issue. "Temporary" without a plan is permanent.
16
+ - **Does the on path get the normal review?** Lower polish bar, not zero. Security and data-integrity questions still apply — an experiment that corrupts data is still corruption.
17
+ - **Public surface.** Does the flag itself become public API? If users can set it, removing it later is a breaking change.
18
+ - **Tests.** Both paths tested? The off path especially — it's what everyone runs.
19
+ - **Docs.** Is it documented as experimental, or documented as if stable?
20
+
21
+ ## Signals → branches
22
+
23
+ - Flag read in more than ~3 places → request consolidation so the flag can be removed cleanly
24
+ - Off path has new imports / side effects at module load → leak; request isolation before merge
25
+ - Flag exposed in public config with no "experimental" marking → it's public API now; changeset and docs questions apply
26
+ - No tracking issue / no owner → request one, and ask for the removal plan
27
+ - Experiment touches storage or wire format → load `schema-storage.md` regardless of the flag; data written under the experiment outlives it
28
+
29
+ ## Verify
30
+
31
+ - **Run the off path** on the branch with the flag unset; diff observable behavior against base. It should be identical; record any meaningful difference.
32
+ - Grep every read of the flag: `grep -rn "<flag name>"` on the branch — count and list.
33
+ - Confirm the default with a pointer.
34
+ - Tests: `grep -n "<flag>" <test files>` — both branches exercised?
@@ -0,0 +1,37 @@
1
+ # Infra / build / CI / tooling
2
+
3
+ Build config, CI workflows, lint/format config, scripts, dev tooling, monorepo plumbing, release machinery.
4
+
5
+ **Reviewing:** blast radius on every developer and every future PR.
6
+ **Read first:** what happens to someone who pulls this tomorrow — a fresh clone, a stale branch, a rebase.
7
+ **Done means:** testable before merge; rollback is clear; the change can't break `main` for everyone on landing.
8
+ **Trap:** works on the author's machine.
9
+ **Attention:** the failure modes on machines that aren't the author's — CI runners, fresh clones, Windows, older Node, no network.
10
+
11
+ ## Questions
12
+
13
+ - **Who does this hit?** Every developer, every CI run, every release? Or one package's tests? Blast radius sets the depth.
14
+ - **Can it be tested before merge?** CI changes are notoriously only testable by merging. Did the author run it on a branch? Is there output?
15
+ - **What's the rollback?** A revert is enough — unless the change also mutated caches, lockfiles, generated files, or published something.
16
+ - **Fresh clone.** Does `pnpm install && pnpm build` still work from nothing? Does a stale branch rebased onto this break in a way the error message explains?
17
+ - **Lockfile changes.** Are they explained? A lockfile diff with no dependency change in `package.json` is a question.
18
+ - **Determinism and platform.** Path separators, shell (`sh` vs `bash` vs `zsh`), `timeout` doesn't exist on macOS, GNU vs BSD flags, case sensitivity, line endings.
19
+ - **Secrets and permissions.** CI workflows that gain a token or a permission scope. Least privilege?
20
+ - **Speed.** Does this add minutes to every CI run? Cache behavior?
21
+ - **Does it change what gets published?** Build output, exports map, `files` in package.json — the most silent breakage there is.
22
+ - **Is it the documented shape?** AGENTS.md / CONTRIBUTING describe how to build and test. Does this change keep them true, or do the docs need to change with it?
23
+
24
+ ## Signals → branches
25
+
26
+ - Workflow file changed with no evidence it ran → ask for the run link; without it, `inferred`
27
+ - Lockfile changed alone → find the cause; unexplained lockfile churn is a finding
28
+ - Build output or exports map touched → this is a packaging change; verify with `pnpm pack` and a consumer install
29
+ - Shell script uses bash-isms or GNU flags → platform finding
30
+ - Cache key changed → check both the hit path and the miss path
31
+
32
+ ## Verify
33
+
34
+ - **Fresh-clone install and build** in a temp worktree when install/build behavior is part of the claim: use the repository's documented command shape and record the decisive result.
35
+ - If a workflow changed, find the run for this branch in `gh run list --branch <head>` and read it.
36
+ - If packaging changed: `pnpm pack` the affected package, install the tarball into a scratch project, import the public entry.
37
+ - Run the script on macOS if the author is on Linux, or note that you couldn't.
@@ -0,0 +1,35 @@
1
+ # New internal capability (no public surface)
2
+
3
+ **Reviewing:** architectural fit.
4
+ **Read first:** the closest sibling feature, side by side with the new one.
5
+ **Done means:** uses the same primitives and layers as the sibling; hooks into what it should; nothing reinvented.
6
+ **Trap:** reinvented wheel; bypassed layer.
7
+ **Attention:** the seams — where the new code touches the existing system. The internals of the new code matter less than what it goes through and around.
8
+
9
+ ## Questions
10
+
11
+ - **Find the sibling.** Whatever this does, something similar exists. Read how _it_ does it. Same primitives, or a second event bus / retry loop / cache / helper that exists three directories over? Reinvention is the #1 sign of not knowing the internals.
12
+ - **Trace one call end-to-end.** Main entry point to the bottom, every layer. It should go through what everything else goes through (validation, logging, tracing, storage abstraction) rather than around. The bypass is invisible in the diff — it's the _absence_ of a call — so you have to trace, not read.
13
+ - **What should it have hooked into?** Lifecycle hooks, middleware, processors, registries, cleanup. Capability that doesn't register is invisible to observability, plugins, teardown.
14
+ - **Which existing code did the author model this on?** A good answer names a file. No answer means they didn't look.
15
+ - **Should this exist, and here?** Do we want to own this? Does it belong in this package, this layer? Would an example or a plugin have served?
16
+ - **Is it over-built?** Abstraction with one caller, options no caller sets, a registry for two entries, indirection between two things that could talk directly. What does it buy _today_?
17
+ - **What does this assume, and where is it enforced?** "Always called after init," "never concurrent." Name it; find the enforcement.
18
+ - **Resource lifecycle.** Handles, listeners, timers, subscriptions, child processes — everything opened must be closed. Leaks don't show in tests.
19
+ - **What will the next PR in this area need?** Does this make it easier or harder? Abstractions that fit one use case make the second worse.
20
+ - **Tests as spec.** Cover the implementation, read only the tests — can you reconstruct what it does? Tests of implementation details ("calls internal helper with args") break on refactor and catch no bugs.
21
+
22
+ ## Signals → branches
23
+
24
+ - A helper/utility added that has a near-twin in the repo → grep for the twin; request consolidation; two implementations of one thing will diverge
25
+ - Entry point doesn't pass through the layer the sibling does → request routing through the layer before merge unless the author can justify the bypass
26
+ - No sibling exists → the "should this exist" question gets more weight, and the design review is on you — compare against the three closest things anyway
27
+ - "Extensible," "pluggable," "generic" in the description → over-built check
28
+ - Opens something (listener, timer, stream) with no matching close → trace teardown
29
+
30
+ ## Verify
31
+
32
+ - Compare the sibling's entry point with the new one; record only divergences that affect the review.
33
+ - Trace the layers the sibling's call passes through and confirm the important ones on the new path with pointers.
34
+ - Grep for near-twin helpers: `archaeology.md` → Callers / twin search.
35
+ - If it's exercisable without a public surface (internal test harness, existing example), run it.
@@ -0,0 +1,35 @@
1
+ # Large / cross-cutting
2
+
3
+ Many packages, many concerns, or simply too big to review in one sitting.
4
+
5
+ **Reviewing:** whether it should be one PR at all.
6
+ **Read first:** the structure — the file list grouped by package and by kind of change — not the content.
7
+ **Done means:** either split into reviewable units, or reviewed as a series with a design doc or plan up front, each unit under its own category's rules.
8
+ **Trap:** reviewing 3,000 lines in one sitting and calling it done.
9
+ **Attention:** the structure and the seams between concerns. Content review happens per-unit, after the structure question is settled.
10
+
11
+ ## Questions
12
+
13
+ - **How many independent changes are in here?** Group the file list by concern. Each group that could land on its own is a candidate PR. Independent changes in one PR get one review's attention divided by N.
14
+ - **Can it be reviewed in one sitting?** If you can't, you're not reviewing it — you're skimming it. That's the honest answer, and it's the finding.
15
+ - **Is there a plan or design doc?** Large changes without a stated design are a design review disguised as a code review. Ask for the design; review that first.
16
+ - **What's the merge risk?** How long has this been open; how far has the base moved; how many in-flight PRs touch the same files.
17
+ - **Which parts are which category?** A large PR is always mixed. Triage each portion: the refactor part, the API part, the fix part. Each gets its own page and its own bar.
18
+ - **Is the size the problem, or the scope?** 2,000 lines of generated code is mechanical. 2,000 lines of hand-written changes across five packages is scope.
19
+ - **Can it be reviewed as a stack?** If the author can't split, can the commits be reviewed in order as if they were a stack? Only if the commits are clean units.
20
+ - **What does "approve" mean here?** You cannot vouch for 3,000 lines. Be explicit about what you reviewed deeply and what you skimmed — in the handoff and, if posted, in the review comment.
21
+
22
+ ## Signals → branches
23
+
24
+ - File list spans >2 packages with different kinds of change → "should be split" finding; propose the split
25
+ - No design doc and the description is a change list → needs discussion; ask for the design before the code
26
+ - Commits are "wip," "fix," "more" → cannot be reviewed as a stack; split is the only path
27
+ - Commits are clean units → review in commit order; each commit under its own category
28
+
29
+ ## Verify
30
+
31
+ - Group `git diff <base>...HEAD --stat` by concern and propose a split when it would improve reviewability.
32
+ - Use `git log --oneline <base>..HEAD` to assess whether commits are reviewable units.
33
+ - Verify high-risk units under their relevant category guidance.
34
+ - Search for in-flight PRs touching the same paths when merge-order risk is plausible.
35
+ - Be explicit about what you reviewed deeply and what you did not; no arbitrary line threshold determines sufficiency.
@@ -0,0 +1,36 @@
1
+ # Mechanical
2
+
3
+ Typo, formatting, lockfile, version bump, dependency bump, generated code, rename-only.
4
+
5
+ **Reviewing:** that it's _only_ that.
6
+ **Read first:** the file list and the line count. Not the content.
7
+ **Done means:** CI green; nothing else in the diff; the generated output matches the generator.
8
+ **Trap:** something real hiding in 2,000 lines of generated noise.
9
+ **Attention:** the parts that aren't mechanical. Find them, or confirm there are none.
10
+
11
+ ## Questions
12
+
13
+ - **Is every hunk the same kind of change?** A formatting PR with one hunk that isn't formatting. A rename with one hunk that changes a default. Scan for the odd one out; that's the whole review.
14
+ - **Line count vs. what the change should need.** A version bump that touches 40 files — why? A typo fix that touches a lockfile — why?
15
+ - **Generated code: does it match the generator?** Regenerate and diff. Hand edits to generated files are a finding.
16
+ - **Dependency bump: what changed in the dependency?** Read the dependency's changelog between the two versions. A "patch" bump can carry a behavior change. A major bump is not mechanical — re-categorize.
17
+ - **Lockfile: is the diff explained by the `package.json` diff?** Unexplained lockfile churn is a question.
18
+ - **Rename: every caller?** A rename that misses a string reference (config key, docs, error message, a dynamic import) breaks silently.
19
+ - **Does CI actually cover it?** Formatting changes in files CI doesn't lint; generated files CI doesn't regenerate.
20
+
21
+ ## Signals → branches
22
+
23
+ - One hunk that isn't the same kind as the rest → re-categorize that hunk under its real category; the mixing is itself a finding (it hid a real change in noise)
24
+ - Dependency major bump → not mechanical; load `behavior-change.md` and read the dependency's migration guide
25
+ - Generated file with no generator run in the commit list → regenerate and diff
26
+ - Rename with fewer callers updated than the grep finds → request the missing callers before merge
27
+
28
+ ## Verify
29
+
30
+ - Categorize every hunk by kind: `git diff <base>...HEAD --stat` then skim each file's diff for the odd hunk.
31
+ - Generated code: run the generator on the branch; `git status` should be clean.
32
+ - Dependency bump: `gh api repos/<dep-owner>/<dep-repo>/releases` or the package's CHANGELOG for the range; note anything behavioral.
33
+ - Rename: grep the old name repo-wide on the branch, including docs, config, and string literals.
34
+ - CI green: `gh pr checks`.
35
+
36
+ Most mechanical PRs produce a handoff of Triage + Verdict. That's correct, not lazy.
@@ -0,0 +1,33 @@
1
+ # Performance
2
+
3
+ **Reviewing:** numbers.
4
+ **Read first:** the benchmark and its input shape — before the code change.
5
+ **Done means:** before/after measured on a representative input; behavior unchanged; the win is where the description says it is.
6
+ **Trap:** faster on a toy input; correctness quietly changed to get the number.
7
+ **Attention:** the measurement first, then equivalence. A perf PR is a refactor with a number attached — it has to pass the refactor bar _and_ show the number.
8
+
9
+ ## Questions
10
+
11
+ - **What's n in production?** Not "is it O(n)" — what's the real input size and shape? A win at 50 items that regresses at 50k, or vice versa, is a loss.
12
+ - **Is this the hot path?** An extra allocation per token matters; in setup it doesn't. Conversely, an optimization off the hot path is complexity for nothing.
13
+ - **What was measured, how, on what?** Wall clock? Allocations? Under load? Cold or warm? One run or many? Was the baseline measured the same way, on the same machine, same session?
14
+ - **Did behavior change?** Perf PRs remove work. Was any of that work load-bearing — an ordering guarantee, a validation, a flush? Apply the refactor questions: tests unchanged, deletions justified.
15
+ - **Is the complexity worth the number?** A 3% win that adds a cache with invalidation logic is a net loss in maintenance. Ask what the number _buys_ a user.
16
+ - **Does the win survive the real environment?** Node version, bundler, the actual storage backend, network in the loop.
17
+ - **Caches and memoization.** Invalidation, memory growth, key collisions, behavior when the cache is cold.
18
+ - **Is there a simpler win?** Sometimes the measured bottleneck is one line away from a trivial fix and the PR built a subsystem around it.
19
+
20
+ ## Signals → branches
21
+
22
+ - No numbers in the description → the claim is unverified; ask for the benchmark or run one
23
+ - Numbers but no input description → run at 10× and 100× the likely input; report both
24
+ - Tests changed → re-categorize as behavior change; the number may have been bought with correctness
25
+ - A cache or memo introduced → invalidation and memory questions; look for the test that exercises invalidation
26
+ - "Should be faster" without measurement → inferred; not a perf PR until measured
27
+
28
+ ## Verify
29
+
30
+ - **Reproduce the measurement** on base and branch with the same harness and input. `hyperfine` works for CLI-shaped things; a scripted loop can measure library calls. Record the method and material result.
31
+ - **Scale the input** to production shape and re-run.
32
+ - Refactor verification: tests untouched (`git diff <base>...HEAD --stat -- '**/*.test.*'` empty), suite green.
33
+ - If a cache was added: write a case that invalidates and confirm behavior.