@prestyj/cli 5.28.0 → 5.29.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (398) hide show
  1. package/README.md +2 -2
  2. package/assets/motion/THIRD-PARTY.md +14 -19
  3. package/assets/motion/bin/contact-sheet.mjs +9 -3
  4. package/assets/motion/bin/cues.mjs +337 -0
  5. package/assets/motion/bin/flash-check.mjs +314 -0
  6. package/assets/motion/bin/fonts.mjs +17 -13
  7. package/assets/motion/bin/library.mjs +61 -18
  8. package/assets/motion/bin/motion-blur.mjs +943 -0
  9. package/assets/motion/bin/motion-check.mjs +366 -9
  10. package/assets/motion/bin/music-fit.mjs +436 -0
  11. package/assets/motion/bin/pdf-extract.mjs +22 -10
  12. package/assets/motion/bin/reference-study.mjs +346 -0
  13. package/assets/motion/bin/score-synth.mjs +1052 -93
  14. package/assets/motion/fonts/finger-paint/OFL.txt +93 -0
  15. package/assets/motion/fonts/finger-paint/finger-paint-normal.woff2 +0 -0
  16. package/assets/motion/fonts/fonts.json +56 -0
  17. package/assets/motion/fonts/short-stack/OFL.txt +94 -0
  18. package/assets/motion/fonts/short-stack/short-stack-normal.woff2 +0 -0
  19. package/assets/motion/fonts/sora/OFL.txt +93 -0
  20. package/assets/motion/fonts/sora/sora-normal.woff2 +0 -0
  21. package/assets/motion/fonts/specimen.jpg +0 -0
  22. package/assets/motion/fonts/unbounded/OFL.txt +93 -0
  23. package/assets/motion/fonts/unbounded/unbounded-normal.woff2 +0 -0
  24. package/assets/motion/library/README.md +52 -14
  25. package/assets/motion/library/kit/moves.js +1981 -0
  26. package/assets/motion/library/library.json +232 -0
  27. package/assets/motion/library/pieces/camera-rig/meta.json +13 -0
  28. package/assets/motion/library/pieces/camera-rig/piece.html +153 -0
  29. package/assets/motion/library/pieces/camera-rig/preview.jpg +0 -0
  30. package/assets/motion/library/pieces/chain-knock/meta.json +13 -0
  31. package/assets/motion/library/pieces/chain-knock/piece.html +195 -0
  32. package/assets/motion/library/pieces/chain-knock/preview.jpg +0 -0
  33. package/assets/motion/library/pieces/gather-to-logo/meta.json +13 -0
  34. package/assets/motion/library/pieces/gather-to-logo/piece.html +159 -0
  35. package/assets/motion/library/pieces/gather-to-logo/preview.jpg +0 -0
  36. package/assets/motion/library/pieces/morph-carry/meta.json +13 -0
  37. package/assets/motion/library/pieces/morph-carry/piece.html +173 -0
  38. package/assets/motion/library/pieces/morph-carry/preview.jpg +0 -0
  39. package/assets/motion/library/pieces/one-shape-journey/meta.json +13 -0
  40. package/assets/motion/library/pieces/one-shape-journey/piece.html +195 -0
  41. package/assets/motion/library/pieces/one-shape-journey/preview.jpg +0 -0
  42. package/assets/motion/library/pieces/open-from-subject/meta.json +13 -0
  43. package/assets/motion/library/pieces/open-from-subject/piece.html +168 -0
  44. package/assets/motion/library/pieces/open-from-subject/preview.jpg +0 -0
  45. package/assets/motion/library/pieces/request-to-result/meta.json +13 -0
  46. package/assets/motion/library/pieces/request-to-result/piece.html +212 -0
  47. package/assets/motion/library/pieces/request-to-result/preview.jpg +0 -0
  48. package/assets/motion/library/pieces/scale-dive/meta.json +13 -0
  49. package/assets/motion/library/pieces/scale-dive/piece.html +321 -0
  50. package/assets/motion/library/pieces/scale-dive/preview.jpg +0 -0
  51. package/assets/motion/library/pieces/screen-replica-steps/meta.json +13 -0
  52. package/assets/motion/library/pieces/screen-replica-steps/piece.html +366 -0
  53. package/assets/motion/library/pieces/screen-replica-steps/preview.jpg +0 -0
  54. package/assets/motion/library/pieces/zoom-into-card/meta.json +13 -0
  55. package/assets/motion/library/pieces/zoom-into-card/piece.html +179 -0
  56. package/assets/motion/library/pieces/zoom-into-card/preview.jpg +0 -0
  57. package/assets/motion/library/sheets/diagram.jpg +0 -0
  58. package/assets/motion/library/sheets/frame.jpg +0 -0
  59. package/assets/motion/library/sheets/transition.jpg +0 -0
  60. package/assets/motion/library/sheets/ui.jpg +0 -0
  61. package/assets/motion/plugin.json +1 -1
  62. package/assets/motion/references/build-sheet.md +206 -0
  63. package/assets/motion/references/runtime/determinism-rules.md +3 -3
  64. package/assets/motion/references/runtime/gsap-easing-and-stagger.md +28 -28
  65. package/assets/motion/references/runtime/gsap.md +3 -3
  66. package/assets/motion/references/runtime/inputs-and-assets.md +21 -16
  67. package/assets/motion/references/runtime/lint-validate-inspect.md +3 -3
  68. package/assets/motion/references/runtime/minimal-composition.md +6 -0
  69. package/assets/motion/references/runtime/preview-render.md +3 -3
  70. package/assets/motion/skills/app-walkthrough/SKILL.md +66 -0
  71. package/assets/motion/skills/before-after/SKILL.md +53 -0
  72. package/assets/motion/skills/brand-kit/SKILL.md +17 -16
  73. package/assets/motion/skills/dev-tool-video/SKILL.md +57 -0
  74. package/assets/motion/skills/launch-video/SKILL.md +62 -0
  75. package/assets/motion/skills/match-reference/SKILL.md +58 -0
  76. package/assets/motion/skills/motion/SKILL.md +167 -98
  77. package/assets/motion/skills/source-ingest/SKILL.md +23 -15
  78. package/assets/motion/skills/website-video/SKILL.md +59 -0
  79. package/assets/skills/bulletproof/SKILL.md +36 -11
  80. package/assets/skills/bulletproof/references/agent-surface.md +19 -9
  81. package/assets/skills/bulletproof/references/audit-protocol.md +20 -5
  82. package/assets/skills/bulletproof/references/platform-playbooks.md +5 -4
  83. package/assets/skills/bulletproof/references/provenance.md +26 -1
  84. package/assets/skills/bulletproof/references/secure-defaults.md +6 -5
  85. package/assets/skills/bulletproof/references/supply-chain.md +21 -17
  86. package/assets/skills/bulletproof/references/threat-landscape.md +28 -26
  87. package/assets/skills/bulletproof/references/verification.md +2 -0
  88. package/assets/skills/clarify/SKILL.md +25 -16
  89. package/assets/skills/code-review/SKILL.md +71 -13
  90. package/assets/skills/code-review/references/agent-diffs.md +27 -0
  91. package/assets/skills/code-review/references/tests.md +19 -0
  92. package/assets/skills/compliance-guard/SKILL.md +20 -5
  93. package/assets/skills/compliance-guard/references/artifacts.md +1 -1
  94. package/assets/skills/compliance-guard/references/eu-uk.md +16 -16
  95. package/assets/skills/compliance-guard/references/lawsuit-vectors.md +5 -5
  96. package/assets/skills/compliance-guard/references/provenance.md +41 -2
  97. package/assets/skills/compliance-guard/references/sector-gates.md +3 -3
  98. package/assets/skills/compliance-guard/references/security-baseline.md +2 -2
  99. package/assets/skills/compliance-guard/references/trigger-map.md +5 -5
  100. package/assets/skills/compliance-guard/references/us.md +27 -21
  101. package/assets/skills/durable/SKILL.md +87 -79
  102. package/assets/skills/durable/references/agent-db-safety.md +69 -0
  103. package/assets/skills/durable/references/backups-and-runtime.md +19 -12
  104. package/assets/skills/durable/references/migrations-and-schema.md +13 -6
  105. package/assets/skills/evidence-led-ui/SKILL.md +69 -127
  106. package/assets/skills/evidence-led-ui/references/anti-defaults.md +107 -208
  107. package/assets/skills/evidence-led-ui/references/direction.md +124 -0
  108. package/assets/skills/evidence-led-ui/references/production-contract.md +8 -0
  109. package/assets/skills/evidence-led-ui/references/provenance.md +24 -1
  110. package/assets/skills/lean/SKILL.md +90 -71
  111. package/assets/skills/lean/references/memory-and-processes.md +3 -2
  112. package/assets/skills/lean/references/playbooks.md +37 -12
  113. package/assets/skills/refactoring/SKILL.md +24 -3
  114. package/assets/skills/refactoring/references/agent-pitfalls.md +4 -1
  115. package/assets/skills/refactoring/references/legacy.md +21 -0
  116. package/assets/skills/root-cause/SKILL.md +20 -10
  117. package/assets/skills/shared-language/SKILL.md +16 -14
  118. package/assets/skills/tdd/SKILL.md +27 -15
  119. package/dist/app-sidecar.js +203 -47
  120. package/dist/app-sidecar.js.map +1 -1
  121. package/dist/cli.js +20 -28
  122. package/dist/cli.js.map +1 -1
  123. package/dist/core/acceptance-checks.d.ts +48 -0
  124. package/dist/core/acceptance-checks.js +144 -0
  125. package/dist/core/acceptance-checks.js.map +1 -0
  126. package/dist/core/agent-session.d.ts +119 -75
  127. package/dist/core/agent-session.js +561 -395
  128. package/dist/core/agent-session.js.map +1 -1
  129. package/dist/core/agents.d.ts +6 -5
  130. package/dist/core/agents.js +12 -3
  131. package/dist/core/agents.js.map +1 -1
  132. package/dist/core/ask-user.d.ts +90 -8
  133. package/dist/core/ask-user.js +124 -13
  134. package/dist/core/ask-user.js.map +1 -1
  135. package/dist/core/auth-providers.js +1 -1
  136. package/dist/core/auth-providers.js.map +1 -1
  137. package/dist/core/autopilot-verdict.d.ts +5 -1
  138. package/dist/core/autopilot-verdict.js +30 -13
  139. package/dist/core/autopilot-verdict.js.map +1 -1
  140. package/dist/core/bundled-agents.js +1 -3
  141. package/dist/core/bundled-agents.js.map +1 -1
  142. package/dist/core/cache-diagnostics.d.ts +68 -0
  143. package/dist/core/cache-diagnostics.js +196 -0
  144. package/dist/core/cache-diagnostics.js.map +1 -0
  145. package/dist/core/cache-expiry.d.ts +87 -0
  146. package/dist/core/cache-expiry.js +111 -0
  147. package/dist/core/cache-expiry.js.map +1 -0
  148. package/dist/core/compaction/compactor.js +78 -48
  149. package/dist/core/compaction/compactor.js.map +1 -1
  150. package/dist/core/compaction/plan-step-policy.d.ts +46 -0
  151. package/dist/core/compaction/plan-step-policy.js +57 -0
  152. package/dist/core/compaction/plan-step-policy.js.map +1 -0
  153. package/dist/core/destructive-git-guard.d.ts +90 -0
  154. package/dist/core/destructive-git-guard.js +871 -0
  155. package/dist/core/destructive-git-guard.js.map +1 -0
  156. package/dist/core/event-bus.d.ts +3 -0
  157. package/dist/core/event-bus.js +5 -0
  158. package/dist/core/event-bus.js.map +1 -1
  159. package/dist/core/fast-apply-benchmark.d.ts +1 -1
  160. package/dist/core/fast-apply-benchmark.js +2 -2
  161. package/dist/core/fast-apply-benchmark.js.map +1 -1
  162. package/dist/core/injection-detect.d.ts +38 -0
  163. package/dist/core/injection-detect.js +232 -0
  164. package/dist/core/injection-detect.js.map +1 -0
  165. package/dist/core/keep-awake.d.ts +88 -0
  166. package/dist/core/keep-awake.js +251 -0
  167. package/dist/core/keep-awake.js.map +1 -0
  168. package/dist/core/mcp/client.d.ts +72 -0
  169. package/dist/core/mcp/client.js +264 -41
  170. package/dist/core/mcp/client.js.map +1 -1
  171. package/dist/core/mcp/content.js +6 -2
  172. package/dist/core/mcp/content.js.map +1 -1
  173. package/dist/core/mcp/store.d.ts +6 -1
  174. package/dist/core/mcp/store.js +12 -1
  175. package/dist/core/mcp/store.js.map +1 -1
  176. package/dist/core/mcp/types.d.ts +18 -0
  177. package/dist/core/model-unavailable.d.ts +14 -0
  178. package/dist/core/model-unavailable.js +23 -0
  179. package/dist/core/model-unavailable.js.map +1 -0
  180. package/dist/core/node-debugger.d.ts +148 -0
  181. package/dist/core/node-debugger.js +642 -0
  182. package/dist/core/node-debugger.js.map +1 -0
  183. package/dist/core/nolan-context.d.ts +7 -5
  184. package/dist/core/nolan-context.js +106 -16
  185. package/dist/core/nolan-context.js.map +1 -1
  186. package/dist/core/nolan-prompt.js +24 -21
  187. package/dist/core/nolan-prompt.js.map +1 -1
  188. package/dist/core/package-threats.d.ts +18 -0
  189. package/dist/core/package-threats.js +168 -0
  190. package/dist/core/package-threats.js.map +1 -0
  191. package/dist/core/persistent-shell.d.ts +58 -6
  192. package/dist/core/persistent-shell.js +331 -49
  193. package/dist/core/persistent-shell.js.map +1 -1
  194. package/dist/core/process-manager.d.ts +14 -0
  195. package/dist/core/process-manager.js +61 -0
  196. package/dist/core/process-manager.js.map +1 -1
  197. package/dist/core/progress/git-xp.js +8 -14
  198. package/dist/core/progress/git-xp.js.map +1 -1
  199. package/dist/core/project-discovery.js +77 -26
  200. package/dist/core/project-discovery.js.map +1 -1
  201. package/dist/core/semantic-search-benchmark.d.ts +1 -1
  202. package/dist/core/semantic-search-benchmark.js +2 -2
  203. package/dist/core/semantic-search-benchmark.js.map +1 -1
  204. package/dist/core/session-history.d.ts +12 -0
  205. package/dist/core/session-history.js +27 -0
  206. package/dist/core/session-history.js.map +1 -1
  207. package/dist/core/session-manager.d.ts +13 -1
  208. package/dist/core/session-manager.js +38 -18
  209. package/dist/core/session-manager.js.map +1 -1
  210. package/dist/core/session-summary-index.d.ts +37 -0
  211. package/dist/core/session-summary-index.js +172 -0
  212. package/dist/core/session-summary-index.js.map +1 -0
  213. package/dist/core/settings-manager.d.ts +2 -0
  214. package/dist/core/settings-manager.js +10 -0
  215. package/dist/core/settings-manager.js.map +1 -1
  216. package/dist/core/shell-threats-popular-packages.d.ts +11 -0
  217. package/dist/core/shell-threats-popular-packages.js +675 -0
  218. package/dist/core/shell-threats-popular-packages.js.map +1 -0
  219. package/dist/core/shell-threats.d.ts +8 -0
  220. package/dist/core/shell-threats.js +186 -0
  221. package/dist/core/shell-threats.js.map +1 -0
  222. package/dist/core/skills.js +16 -4
  223. package/dist/core/skills.js.map +1 -1
  224. package/dist/core/stream-rules.d.ts +30 -0
  225. package/dist/core/stream-rules.js +151 -0
  226. package/dist/core/stream-rules.js.map +1 -0
  227. package/dist/core/subagent-manager.d.ts +23 -6
  228. package/dist/core/subagent-manager.js +25 -9
  229. package/dist/core/subagent-manager.js.map +1 -1
  230. package/dist/core/subagent-policy.js +1 -1
  231. package/dist/core/subagent-policy.js.map +1 -1
  232. package/dist/core/subagent-receipt.d.ts +54 -0
  233. package/dist/core/subagent-receipt.js +276 -0
  234. package/dist/core/subagent-receipt.js.map +1 -0
  235. package/dist/core/subagent-turn-record.d.ts +2 -0
  236. package/dist/core/subagent-turn-record.js.map +1 -1
  237. package/dist/core/test-impact.d.ts +73 -0
  238. package/dist/core/test-impact.js +467 -0
  239. package/dist/core/test-impact.js.map +1 -0
  240. package/dist/core/thinking-level.d.ts +1 -1
  241. package/dist/core/thinking-level.js +1 -1
  242. package/dist/core/thinking-level.js.map +1 -1
  243. package/dist/core/verification-gate.d.ts +2 -0
  244. package/dist/core/verification-gate.js +4 -0
  245. package/dist/core/verification-gate.js.map +1 -1
  246. package/dist/core/verification-snapshot.js +3 -5
  247. package/dist/core/verification-snapshot.js.map +1 -1
  248. package/dist/core/workspace-guard.d.ts +18 -7
  249. package/dist/core/workspace-guard.js +227 -60
  250. package/dist/core/workspace-guard.js.map +1 -1
  251. package/dist/core/worktree-setup.d.ts +23 -0
  252. package/dist/core/worktree-setup.js +128 -9
  253. package/dist/core/worktree-setup.js.map +1 -1
  254. package/dist/core/worktree.d.ts +20 -2
  255. package/dist/core/worktree.js +74 -31
  256. package/dist/core/worktree.js.map +1 -1
  257. package/dist/interactive.js +2 -1
  258. package/dist/interactive.js.map +1 -1
  259. package/dist/modes/json-mode.js +11 -2
  260. package/dist/modes/json-mode.js.map +1 -1
  261. package/dist/modes/subagent-worker-mode.d.ts +36 -1
  262. package/dist/modes/subagent-worker-mode.js +100 -42
  263. package/dist/modes/subagent-worker-mode.js.map +1 -1
  264. package/dist/motion-agent/motion-agent.d.ts +7 -3
  265. package/dist/motion-agent/motion-agent.js +6 -8
  266. package/dist/motion-agent/motion-agent.js.map +1 -1
  267. package/dist/motion-agent/motion-prompt.d.ts +1 -1
  268. package/dist/motion-agent/motion-prompt.js +16 -25
  269. package/dist/motion-agent/motion-prompt.js.map +1 -1
  270. package/dist/motion-agent/motion-review-session.js +1 -1
  271. package/dist/motion-agent/motion-review-session.js.map +1 -1
  272. package/dist/motion-agent/motion-review.d.ts +10 -3
  273. package/dist/motion-agent/motion-review.js +15 -7
  274. package/dist/motion-agent/motion-review.js.map +1 -1
  275. package/dist/motion-agent/motion-studio-context.js +1 -1
  276. package/dist/motion-agent/motion-studio-context.js.map +1 -1
  277. package/dist/system-prompt.d.ts +2 -1
  278. package/dist/system-prompt.js +15 -4
  279. package/dist/system-prompt.js.map +1 -1
  280. package/dist/test-support/keep-alive.d.ts +14 -0
  281. package/dist/test-support/keep-alive.js +17 -0
  282. package/dist/test-support/keep-alive.js.map +1 -0
  283. package/dist/tools/ask-user.js +3 -3
  284. package/dist/tools/ask-user.js.map +1 -1
  285. package/dist/tools/bash-read-evidence.d.ts +10 -0
  286. package/dist/tools/bash-read-evidence.js +133 -0
  287. package/dist/tools/bash-read-evidence.js.map +1 -0
  288. package/dist/tools/bash.d.ts +10 -1
  289. package/dist/tools/bash.js +179 -7
  290. package/dist/tools/bash.js.map +1 -1
  291. package/dist/tools/debug.d.ts +54 -0
  292. package/dist/tools/debug.js +233 -0
  293. package/dist/tools/debug.js.map +1 -0
  294. package/dist/tools/edit.js +13 -5
  295. package/dist/tools/edit.js.map +1 -1
  296. package/dist/tools/goals.d.ts +1 -1
  297. package/dist/tools/index.d.ts +25 -2
  298. package/dist/tools/index.js +48 -7
  299. package/dist/tools/index.js.map +1 -1
  300. package/dist/tools/prompt-hints.js +2 -0
  301. package/dist/tools/prompt-hints.js.map +1 -1
  302. package/dist/tools/read-tracker.d.ts +35 -2
  303. package/dist/tools/read-tracker.js +108 -11
  304. package/dist/tools/read-tracker.js.map +1 -1
  305. package/dist/tools/read.js +39 -7
  306. package/dist/tools/read.js.map +1 -1
  307. package/dist/tools/skill.js +5 -0
  308. package/dist/tools/skill.js.map +1 -1
  309. package/dist/tools/subagent-control.js +44 -8
  310. package/dist/tools/subagent-control.js.map +1 -1
  311. package/dist/tools/subagent-shared.d.ts +48 -8
  312. package/dist/tools/subagent-shared.js +75 -14
  313. package/dist/tools/subagent-shared.js.map +1 -1
  314. package/dist/tools/subagent.d.ts +8 -2
  315. package/dist/tools/subagent.js +28 -10
  316. package/dist/tools/subagent.js.map +1 -1
  317. package/dist/tools/task-output.js +3 -2
  318. package/dist/tools/task-output.js.map +1 -1
  319. package/dist/tools/task-send.d.ts +1 -1
  320. package/dist/tools/task-send.js +15 -1
  321. package/dist/tools/task-send.js.map +1 -1
  322. package/dist/tools/tool-tiers.d.ts +2 -2
  323. package/dist/tools/tool-tiers.js +3 -2
  324. package/dist/tools/tool-tiers.js.map +1 -1
  325. package/dist/tools/truncate.d.ts +21 -0
  326. package/dist/tools/truncate.js +187 -0
  327. package/dist/tools/truncate.js.map +1 -1
  328. package/dist/tools/ui-adopt.js +2 -0
  329. package/dist/tools/ui-adopt.js.map +1 -1
  330. package/dist/tools/write.js +4 -3
  331. package/dist/tools/write.js.map +1 -1
  332. package/dist/ui/App.d.ts +0 -4
  333. package/dist/ui/App.js +5 -28
  334. package/dist/ui/App.js.map +1 -1
  335. package/dist/ui/components/ActivityIndicator.js +1 -0
  336. package/dist/ui/components/ActivityIndicator.js.map +1 -1
  337. package/dist/ui/components/Footer.js +1 -1
  338. package/dist/ui/components/Footer.js.map +1 -1
  339. package/dist/ui/hooks/useAgentLoop.d.ts +1 -8
  340. package/dist/ui/hooks/useAgentLoop.js +1 -119
  341. package/dist/ui/hooks/useAgentLoop.js.map +1 -1
  342. package/dist/ui/render.d.ts +2 -4
  343. package/dist/ui/render.js +3 -2
  344. package/dist/ui/render.js.map +1 -1
  345. package/dist/utils/git.d.ts +77 -0
  346. package/dist/utils/git.js +285 -21
  347. package/dist/utils/git.js.map +1 -1
  348. package/dist/utils/github-ci.js +2 -1
  349. package/dist/utils/github-ci.js.map +1 -1
  350. package/dist/utils/github.js +11 -9
  351. package/dist/utils/github.js.map +1 -1
  352. package/dist/utils/image.d.ts +14 -0
  353. package/dist/utils/image.js +16 -0
  354. package/dist/utils/image.js.map +1 -1
  355. package/dist/utils/process.d.ts +20 -0
  356. package/dist/utils/process.js +98 -0
  357. package/dist/utils/process.js.map +1 -1
  358. package/dist/utils/text.d.ts +12 -0
  359. package/dist/utils/text.js +11 -0
  360. package/dist/utils/text.js.map +1 -1
  361. package/package.json +6 -6
  362. package/assets/motion/skills/mixkit-split-text-617/SKILL.md +0 -115
  363. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-480.json +0 -1
  364. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-494.json +0 -1
  365. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-5.json +0 -1
  366. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-508.json +0 -1
  367. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6525.json +0 -1
  368. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6539.json +0 -1
  369. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6647.json +0 -1
  370. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6663.json +0 -1
  371. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6677.json +0 -1
  372. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6691.json +0 -1
  373. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6706.json +0 -1
  374. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6722.json +0 -1
  375. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6736.json +0 -1
  376. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6750.json +0 -1
  377. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6764.json +0 -1
  378. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6780.json +0 -1
  379. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6794.json +0 -1
  380. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6810.json +0 -1
  381. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6823.json +0 -1
  382. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6872.json +0 -1
  383. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-6936.json +0 -1
  384. package/assets/motion/skills/mixkit-split-text-617/data/compositions/comp-76.json +0 -1
  385. package/assets/motion/skills/mixkit-split-text-617/data/manifest.json +0 -495
  386. package/assets/motion/skills/mixkit-split-text-617/references/RECONSTRUCTION.md +0 -201
  387. package/assets/motion/skills/mixkit-split-text-617/references/VERIFICATION.md +0 -107
  388. package/assets/motion/skills/mixkit-split-text-617/tools/inspect_motion.py +0 -130
  389. package/assets/motion/skills/video-qa/SKILL.md +0 -66
  390. package/dist/core/ideal-review-subagent.d.ts +0 -42
  391. package/dist/core/ideal-review-subagent.js +0 -95
  392. package/dist/core/ideal-review-subagent.js.map +0 -1
  393. package/dist/core/ideal-review.d.ts +0 -82
  394. package/dist/core/ideal-review.js +0 -242
  395. package/dist/core/ideal-review.js.map +0 -1
  396. package/dist/motion-agent/motion-check-tool.d.ts +0 -21
  397. package/dist/motion-agent/motion-check-tool.js +0 -206
  398. package/dist/motion-agent/motion-check-tool.js.map +0 -1
@@ -1,4 +1,4 @@
1
- # Threat Landscape — snapshot 12 August 2026
1
+ # Threat Landscape — snapshot 3 October 2026
2
2
 
3
3
  Why the defaults in this skill are what they are. Confidence markers: **[V]** verified against a primary source at snapshot time, **[S]** snapshot-accurate but volatile, **[U]** uncertain or single-sourced. Preserve the markers when you repeat these claims.
4
4
 
@@ -8,10 +8,9 @@ Read this once per full review to set the threat model. Do not paste incident li
8
8
 
9
9
  **Confirmed: AI-orchestrated intrusion is real and operational.**
10
10
 
11
- - **GTG-1002** [V] — Anthropic disclosed (13 Nov 2025) the first documented largely-autonomous AI-orchestrated espionage campaign: a likely China-nexus actor drove an agent plus tooling through recon, vulnerability discovery, exploitation, lateral movement, credential harvesting and exfiltration against roughly 30 targets, with the large majority of operational tasks machine-executed. Now tracked as MITRE ATT&CK campaign **C0062**.
12
- - **Post-compromise, not just phishing** [V] — Anthropic's ATT&CK mapping of 832 banned accounts (Mar 2025–Mar 2026, published 3 Jun 2026) found AI used most for malware development, with AI-assisted account discovery rising while AI-assisted phishing fell. The durable differentiator is scaffolding that chains stages autonomously, not the operator's skill.
13
- - **Runtime LLM use inside malware** [V] — Google GTIG (5 Nov 2025) documented the first malware families querying a model at runtime: a dropper that requests just-in-time obfuscation, and a data-theft tool with no hard-coded collection commands that prompts a hosted model for them instead. Signature-based detection degrades against code that rewrites itself per execution.
14
- - **Machine-found bugs at scale** [V] — Anthropic's Glasswing programme reported scanning 1,000+ open-source projects, yielding tens of thousands of issues with thousands rated high or critical and a high validation rate on the sampled subset.
11
+ - **GTG-1002** [V] — Anthropic disclosed (13 Nov 2025) the first documented largely-autonomous AI-orchestrated espionage campaign: an actor it assessed as China state-sponsored drove a coding agent wired to tooling over MCP through recon, vulnerability discovery, exploitation, credential harvesting and exfiltration against roughly 30 organizations, with 80–90% of tactical work machine-executed and a small number of intrusions succeeding. The operators got the model to cooperate by **posing as a security firm doing authorized defensive testing** — so a defensive framing in a brief is necessary, never sufficient; the Hard stops still apply.
12
+ - **Runtime LLM use inside malware** [U] — Google GTIG (Nov 2025) reported malware families that query a hosted model at runtime for obfuscation or collection commands. Not re-verified for this snapshot.
13
+ - **Machine-found bugs at scale** [V] — Anthropic's Project Glasswing reported 23,019 vulnerability candidates; by late July 2026 about 126 had become published CVEs and one was confirmed exploited (per VulnCheck). Candidates are not CVEs, and CVEs are not exploitation.
15
14
 
16
15
  **The honest counterweight — do not overstate this.** VulnCheck (28 Jul 2026) [V] found that of ~1,061 vulnerabilities attributed to AI-assisted discovery, only about 1.3% are confirmed exploited in the wild — roughly the same rate as vulnerabilities generally. Discovery volume is not exploitation volume. Machine-scale scanning has moved the bottleneck to **maintainer capacity to triage, patch, test and ship**, which is exactly where a small team is weakest.
17
16
 
@@ -20,50 +19,53 @@ Read this once per full review to set the threat model. Do not paste incident li
20
19
  1. Assume any public repository has been read end-to-end by an automated system. Obscurity was never a control; now it is not even a delay.
21
20
  2. The bug classes machines find fastest are the ones with a cheap verification oracle — web/API classes and memory-safety in parsers. Business logic and multi-actor authorization remain comparatively hard for them, and remain where the expensive breaches happen.
22
21
  3. Patch latency is now the dominant controllable variable. A dependency you cannot update quickly is a standing liability.
23
- 4. Breakout speed is measured in minutes [S] — CrowdStrike's 2026 report cites an average eCrime breakout time under half an hour, with the fastest well under a minute. Detection that requires a human to read a dashboard within the hour is not a control.
22
+ 4. Breakout speed is measured in minutes [U] (vendor threat reports; not re-verified). Detection that requires a human to read a dashboard within the hour is not a control.
24
23
 
25
24
  ## 2. Supply chain — the dominant compromise route for small teams
26
25
 
27
26
  Named incidents, with the fingerprint a defender can grep for. All [V] unless marked.
28
27
 
29
- **The self-propagating npm worm lineage.** Shai-Hulud (Sept 2025) established the pattern: steal a publish token, enumerate every package the victim can publish, inject, republish. CISA warned of 500+ compromised packages targeting source-control and cloud credentials; the marker artifact was a workflow file named `shai-hulud-workflow.yml`. The November 2025 wave added destructive behavior; forensics across ~6,943 compromised machines found tens of thousands of unique secrets, and **59% of compromised machines were CI/CD runners rather than laptops** — CI is the real target.
28
+ **The self-propagating npm worm lineage.** Shai-Hulud (Sept 2025) established the pattern: steal a publish token, enumerate every package the victim can publish, inject, republish. The November 2025 wave added destructive behavior. GitGuardian's analysis of 6,943 compromised developer machines found **59% were CI/CD runners rather than laptops** [V] — CI is the real target.
30
29
 
31
30
  Successor waves worth knowing because each broke a different assumption:
32
31
 
33
- - **Mini Shai-Hulud / TanStack (11 May 2026)** — first credential-free initial access: a fork-triggered workflow with write access to the base repository's cache allowed cache poisoning, and publishing rode the registry's OIDC endpoint. Reported [U] that resulting malicious versions carried valid signed provenance at the highest build level. **Provenance proves where an artifact was built, not that the build was honest.** Treat provenance as necessary, not sufficient.
34
- - **Miasma wave (Jun–Jul 2026)** — dozens of packages under a vendor scope, then several release pipelines of a well-known specification project, each reusing an obfuscated install-time stealer.
35
- - **CHAINDROP (4 Aug 2026)** — the most recent and the most instructive. A maintainer compromise trojanized a monorepo with a worm that backdoored every package that maintainer could publish, reaching packages with very large download counts. Its two novel properties matter more than its scale:
36
- - **Persistence outside the registry.** It committed an agent session-start hook in `.claude/settings.json` and an editor `folderOpen` task in `.vscode/tasks.json`, pushed across many branches. Opening the repository in an editor or an agent was enough to execute. Grep any untrusted repo for both before opening it.
37
- - **Credential sweep aimed at AI accounts.** Its collector targeted coding-assistant and model-provider credentials alongside the usual cloud keys. Your model API keys are now first-class loot.
38
- - Reported mitigations that worked: newer npm versions blocking install hooks by default, eliminating automation tokens that bypass 2FA, and a soak period before adopting new versions.
32
+ - **TanStack / "Mini Shai-Hulud" (11 May 2026)** [V] — 84 malicious versions across 42 `@tanstack/*` packages in about six minutes, with **no npm token stolen**. Chain (from the maintainers' postmortem): a fork PR triggered a `pull_request_target` workflow that ran fork code; it poisoned the shared Actions cache (pnpm store); the release workflow later restored that cache; attacker code then read the OIDC token from runner memory and published. Fixes the maintainers shipped: restructured the PR workflow, repository-owner guards, third-party actions pinned to SHAs, caches purged.
33
+ - **Miasma "Phantom Gyp" waves (mid-2026)** [S] — ran code at install time through a bare `binding.gyp` (implicit `node-gyp rebuild`) with no lifecycle script in `package.json`, evading scanners that only watch `preinstall`/`postinstall`.
34
+ - **ChainDrop (4 Aug 2026)** [V] — a self-propagating worm starting from a malicious commit in `keyv`, reaching hundreds of packages within hours; releases published through hijacked GitHub Actions OIDC carried **valid provenance**. Two properties matter more than scale:
35
+ - **Persistence outside the registry.** Using stolen GitHub tokens it committed a `SessionStart` hook in `.claude/settings.json` and a `runOn: folderOpen` task in `.vscode/tasks.json` to every eligible branch, authored as `claude <claude@users.noreply.github.com>`. Opening the repo in Claude Code or VS Code executed it; no `npm install` needed. Lockfile fixes do not remove it — check all branches.
36
+ - **AI credentials are loot.** It read `.claude`, `.cursor`, and model-provider auth files alongside cloud keys [S].
39
37
 
40
- **CI/CD.** The `tj-actions/changed-files` compromise (Mar 2025, CVE-2025-30066) [V] retroactively repointed version tags at a malicious commit and dumped secrets into build logs across tens of thousands of repositories; a related action compromise enabled it. A 2026 analysis [S] found roughly 38% of organizations still have at least one workflow vulnerable to script injection or a dangerous trigger. In March 2026 a scanner vendor's own action had nearly all of its version tags force-pushed to malicious code [V]. **A mutable tag is not a pin. Pin actions by commit SHA.**
38
+ **Provenance proves where an artifact was built, not that the build was honest.** Both 2026 worms shipped valid provenance. Treat it as necessary, not sufficient.
41
39
 
42
- **Slopsquatting.** Frontier models invent package names at a measurable rate — a 2026 study across ~200,000 prompts found single-digit-percentage hallucination rates, with over a hundred invented names produced identically by every model tested and a large fraction of those names still unregistered at the time of study [S]. Roughly 43% of hallucinated names recur across identical runs, which is what makes them registrable and profitable. **Agents removed the human "does that name look right" checkpoint.** Verify a package exists, is old enough, and is the one you meant, before adding it — see `supply-chain.md`.
40
+ **CI/CD.** The `tj-actions/changed-files` compromise (Mar 2025, CVE-2025-30066) [S] retroactively repointed version tags at a malicious commit and dumped secrets into build logs. **A mutable tag is not a pin. Pin actions by commit SHA** — GitHub itself recommends it and lets admins enforce it [V].
43
41
 
44
- **Editor extensions and MCP servers.** A July–August 2026 campaign published 77 extensions to an open registry that copied real extensions' names and descriptions at version `0.0.1` under namespaces the publishers did not own, beaconing host details on editor start [V]; removal from the registry does not clean already-installed copies. The first malicious MCP server in the wild (Sept 2025) [V] was a clone of a legitimate mail library that added a single silent BCC header — after fifteen clean releases built trust. Separately, hundreds of extension-publisher secrets have leaked, and extensions auto-update by default.
42
+ **Slopsquatting.** Spracklen et al. (USENIX Security 2025) [V]: across 576,000 generated code samples from 16 models, 19.7% of recommended packages did not exist (205,474 unique invented names); commercial models did far better than open ones but not zero. When hallucinating prompts were re-run ten times, 43% of names recurred every time — predictable, so registrable. **Agents removed the human "does that name look right" checkpoint.** Verify a package exists, is old enough, and is the one you meant, before adding it — see `supply-chain.md`.
43
+
44
+ **Editor extensions and MCP servers.** Lookalike-extension campaigns on open registries are recurring [U]; extensions auto-update by default and registry removal does not clean installed copies. The first malicious MCP server in the wild, `postmark-mcp` (Koi Security, Sept 2025) [V], copied the official Postmark server under the same npm name, shipped 15 clean versions, then in 1.0.16 added one line BCC-ing every outgoing email to the attacker.
45
45
 
46
46
  ## 3. AI coding agents as an attack surface
47
47
 
48
48
  If the project you are hardening is itself an agent, a tool server, or ships an AI feature, read `agent-surface.md` in full. The headline facts:
49
49
 
50
- - **Prompt injection has no known reliable prevention** [V]. A 2025 evaluation of twelve proposed defenses reported a 100% bypass rate by adaptive human red-teamers. Design for containment, not for a filter that holds.
50
+ - **Prompt injection has no known reliable prevention** [V]. "The Attacker Moves Second" (Nasr et al., Oct 2025) bypassed 12 published defenses with >90% success for most using adaptive attacks; human red-teaming succeeded on every scenario. Design for containment, not for a filter that holds.
51
51
  - **The lethal trifecta** — private data access + untrusted content + an egress channel. Any two are usually fine; all three is exploitable. This is the single most useful architectural test for an agent feature.
52
- - **Sandbox escapes were the 2026 bumper crop** [S] — multiple critical-severity escapes across the major coding-agent products, including symlink-based escapes and configuration-file protections bypassed from inside the sandbox. The recurring root pattern: **files the agent writes inside the sandbox are later read, loaded, or executed by a trusted process outside it.**
53
- - **Rules-file and skill backdoors** [V] — instructions hidden in `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, or a shared skill using invisible Unicode (tag codepoints, bidi controls, zero-width characters) land directly in the model's context. Documented payloads have instructed agents to exfiltrate local `.env` contents while suppressing output, and to inject credential-harvesting code into every file they generate — turning the agent into the delivery mechanism for a backdoor that reaches CI and production.
54
- - **Fetched content is executable-adjacent** [V] — a malicious issue in a public repository was enough to make an assistant leak private repository contents; a support ticket containing embedded instructions caused an agent holding a privileged database credential to publish secrets back into a public thread. Google reported (Apr 2026) a measurable rise in prompt injections embedded in ordinary web pages [S].
52
+ - **Sandbox escapes were the 2026 bumper crop** [U] — multiple critical-severity escapes across the major coding-agent products, including symlink-based escapes and configuration-file protections bypassed from inside the sandbox. The recurring root pattern: **files the agent writes inside the sandbox are later read, loaded, or executed by a trusted process outside it.**
53
+ - **Rules-file and skill backdoors** [S] — instructions hidden in `CLAUDE.md`, `AGENTS.md`, `.cursorrules`, or a shared skill using invisible Unicode (tag codepoints, bidi controls, zero-width characters) land directly in the model's context. Documented payloads have instructed agents to exfiltrate local `.env` contents while suppressing output, and to inject credential-harvesting code into every file they generate — turning the agent into the delivery mechanism for a backdoor that reaches CI and production.
54
+ - **Fetched content is executable-adjacent** [S] — a malicious issue in a public repository was enough to make an assistant leak private repository contents; a support ticket containing embedded instructions caused an agent holding a privileged database credential to publish secrets back into a public thread.
55
+ - **Repo config is now a worm vector** [V] — ChainDrop (above) persisted in agent and editor config. An agent that auto-loads hooks or auto-runs tasks from the repo it opens is executing the repo author's code.
55
56
 
56
57
  ## 4. What is actually being exploited against small teams
57
58
 
58
- - **Secret sprawl is the number one route.** GitGuardian's 2026 report [V] counted 28.65M new hardcoded secrets on public GitHub in 2025 (+34% year over year), including a large and fast-growing share tied to AI services, thousands of valid credentials inside MCP configuration files, and — the fact that should change behavior — **64% of valid secrets first leaked in 2022 were still unrevoked in 2026**. The same report found a materially higher secret-leak rate in AI-assisted commits than the baseline.
59
- - **Backend-as-a-service row-level security is the highest-yield indie misconfiguration** [S]. A May 2026 study catalogued the failure modes in rank order: RLS disabled entirely; a permissive `using (true)` policy; partial coverage where reads are locked but writes are not; a service-role key shipped in the client bundle; and subtly wrong `auth.uid()` logic. Two details make this worse than it sounds — the dashboard shows an "enabled" badge for the permissive-policy case, and `using (true)` is exactly what a code generator produces when told to "add an RLS policy" without a specific rule. Enumeration is trivial because generated schemas converge on identical table names.
60
- - **Exposed AI and developer infrastructure** [V] — tens of thousands of internet-facing local-inference servers, plus smaller populations of notebook, experiment-tracking and MCP endpoints. Note the honest caveat from the same research: it recorded essentially no AI-aware exploitation; the traffic hitting those ports was generic credential-harvesting scanning probing for `.env` files and cloud secrets. Generic scanners find you first.
61
- - **AI gateways concentrate credentials** [S] — a compromised dependency in a gateway library can expose an organization's entire portfolio of model provider keys at once. Several agent-framework and low-code AI platform CVEs have been used for initial access, credential harvesting and lateral movement, with at least one on CISA's exploited-vulnerabilities catalog.
62
- - **Do not over-rotate to AI, though** [V] — one-third of known-exploited vulnerabilities in the first half of 2026 were content-management systems (largely plugins), with network edge devices generating the rest. If the project runs a CMS or sits behind an appliance, that is the likelier door.
59
+ - **Secret sprawl is the number one route.** GitGuardian's *State of Secrets Sprawl 2026* (Mar 2026) [V]: 28.65M new hardcoded secrets in public GitHub commits in 2025 (+34%); AI-service secrets up 81%; 24,008 unique secrets in public MCP configuration files; Claude Code-assisted commits leaked at 3.2% vs a 1.5% baseline; and **64% of valid secrets from 2022 still valid** — leaks are not revoked.
60
+ - **Backend-as-a-service row-level security is the highest-yield indie misconfiguration.** CVE-2025-48757 [S] covered 170+ apps generated by a vibe-coding platform whose Supabase tables lacked RLS, so the public anon key in the page read and wrote everything. Failure modes to check, all [S]: RLS disabled; a permissive `using (true)` policy (what a generator writes when told "add a policy" with no rule); reads locked but writes open; a service-role key in the client bundle; `auth.uid()` present but not compared to the owner column. Generated schemas converge on the same table names, so enumeration is trivial.
61
+ - **Middleware-only auth** [V] — CVE-2025-29927 let any client skip Next.js middleware (fixed in 15.2.3 / 14.2.25 / 13.5.9) by sending an internal `x-middleware-subrequest` header. Apps that checked auth only in middleware were fully open. Enforce authorization again in the route handler or data layer.
62
+ - **Exposed AI and developer infrastructure** [U] — tens of thousands of internet-facing local-inference servers, plus smaller populations of notebook, experiment-tracking and MCP endpoints. Note the honest caveat from the same research: it recorded essentially no AI-aware exploitation; the traffic hitting those ports was generic credential-harvesting scanning probing for `.env` files and cloud secrets. Generic scanners find you first.
63
+ - **AI gateways concentrate credentials** [U] — a compromised dependency in a gateway library can expose an organization's entire portfolio of model provider keys at once. Several agent-framework and low-code AI platform CVEs have been used for initial access, credential harvesting and lateral movement, with at least one on CISA's exploited-vulnerabilities catalog.
64
+ - **Do not over-rotate to AI, though** [V] — one-third of known-exploited vulnerabilities in the first half of 2026 were content-management systems (VulnCheck 1H-2026), with network edge devices next. If the project runs a CMS or sits behind an appliance, that is the likelier door.
63
65
 
64
66
  ## 5. Speed and economics
65
67
 
66
- - **Median time from CVE publication to confirmed exploitation fell from about 120 days (2025) to about 80 days (1H 2026)** [V]. Roughly a quarter of newly-exploited CVEs showed exploitation on or before publication day. Absolute early-exploitation counts are flat while CVE issuance grew sharply — so the *rate* is falling even as the *speed* rises.
68
+ - **Median time from CVE publication to confirmed exploitation fell from about 120 days (2025) to about 80 days (1H 2026)** [S] (VulnCheck 1H-2026). Known-exploited vulnerabilities grew ~10% while published CVEs grew ~45%, so the exploited *share* is falling even as *speed* rises.
67
69
  - **Leaked credentials are used, not archived.** Assume any secret that touched a public surface, a build log, a paste, or a third-party service is compromised at the moment of exposure. Rotation is the fix; deleting the commit is not.
68
70
  - **Patch aggressively where there is evidence of exploitation.** Guidance in 2026 [S] points toward days, not weeks, for vulnerabilities that are automatable, exploited, and reachable in your deployment.
69
71
 
@@ -9,6 +9,7 @@ Label every finding and every fix with how you know:
9
9
  - **RUNTIME** — you executed something and observed the result. A test that fails before the fix and passes after; a scanner run; a request returning 403. Strongest.
10
10
  - **CODE** — you read the code path end to end and the conclusion follows from what is written. Normal for most review work. Say so.
11
11
  - **DEDUCED** — inferred from framework behavior, convention, or documentation without reading every hop. Acceptable if labelled, never presented as confirmed.
12
+ - **SNAPSHOT** — rests on a dated external fact (a version, default, advisory) from the references. State the date; re-verify before calling it current.
12
13
 
13
14
  Never upgrade a label. "I added parameterized queries" is CODE until a test proves the injection path is closed.
14
15
 
@@ -47,6 +48,7 @@ Pick one per row. Running one scanner in CI beats evaluating five.
47
48
  - **Backend-as-a-service**: query the REST layer directly with the public anon key as an unauthenticated client and as a second user. The dashboard's "enabled" badge is not evidence — a permissive policy shows the same badge [S].
48
49
  - **Mobile**: inspect the built artifact, not the source — extract the bundle and grep for keys; check the manifest's exported components and network security config as they appear in the built app.
49
50
  - **Desktop**: read fuses from the **packaged** application; confirm loopback endpoints reject a request with no token and a wrong `Origin`; confirm the updater rejects an unsigned or downgraded payload.
51
+ - **Supply chain / CI**: `grep -rn 'uses:' .github/workflows | grep -v '@[0-9a-f]\{40\}'` returns nothing; no workflow uses `pull_request_target` with a PR-head checkout; `npm config get allow-scripts`/the allowlist is what you expect; `git log --all -- .claude/settings.json .vscode/tasks.json` shows only commits you made.
50
52
  - **CLI/dev tools**: run against a deliberately hostile fixture repository containing a path-traversal archive entry, a symlink pointing outside the tree, a file name with terminal escape sequences, and a config file with a plugin path. Assert the tool refuses each.
51
53
  - **Contracts**: invariant tests plus a storage-layout diff on every upgradeable deploy.
52
54
  - **ML**: attempt to load a non-safetensors artifact and confirm rejection; confirm inference endpoints are unreachable from outside the private network.
@@ -1,30 +1,39 @@
1
1
  ---
2
2
  name: clarify
3
- description: Use when requirements or a design are genuinely unsettled — the user asks to interrogate, sharpen, or stress-test a plan before building, or mid-build discovery surfaces a decision that materially changes the result. Do NOT use for routine changes, clear bug reports, or work whose requirements are already settled — bias to action there.
3
+ description: Use when requirements or a design are genuinely unsettled — the user asks to interrogate, sharpen, or stress-test a plan or spec before building, a new app or feature request leaves costly-to-reverse product decisions open, or mid-build discovery hits a decision that materially changes the result (scope, data model, UX, tradeoff). Settled terms and hard-to-reverse calls hand off to shared-language. Do NOT use for routine changes, clear bug reports (root-cause), reviewing finished work (code-review), or work whose requirements are already settled — bias to action there.
4
4
  ---
5
5
 
6
6
  # Clarify
7
7
 
8
- Misalignment is the most expensive failure in software: the build succeeds and the thing is still wrong. This skill is a structured interview that settles every open decision **before** implementation. It is not permission-asking — it is decision-forcing.
8
+ Decision-forcing, not permission-asking. Pick the mode:
9
9
 
10
- ## Two modes
10
+ - **Quick gate** (mid-build, a blocking decision appears) → one `ask_user` call with every blocking question; keep building everything that does not depend on the answer. Never halt the whole task at one branch point.
11
+ - **Full interview** (user asks to refine a plan/spec/design) → run the rounds below.
11
12
 
12
- **Quick gate** — mid-build, when you hit an unanswered decision that materially changes the result: batch every blocking question into ONE numbered list, each with your recommended answer, and keep building the parts that don't depend on it. Never stop a whole task at one branch point.
13
+ ## Rules for every question
13
14
 
14
- **Full interview** — the user asks to refine a plan, spec, or design. Run the rounds below.
15
+ 1. **Never ask for facts.** If reading code, running a command, or checking docs can answer it, do that instead. Only decisions reach the user: product calls, taste, tradeoffs with real stakes.
16
+ 2. **Ask only the frontier.** A question is on the frontier when every decision it depends on is settled ("needs persistence?" before "which database?"). Never re-ask a settled decision.
17
+ 3. **One `ask_user` call per round, never prose questions.** Each question: clickable options, your recommended option first and marked, a one-line reason, and the **default** you will apply if skipped ("Default if skipped: SQLite").
18
+ 4. **Show behaviour choices as examples.** When options differ in behaviour, give one concrete case per option ("Given an empty cart, When checkout is pressed, Then …"). Examples expose disagreement that abstract wording hides.
15
19
 
16
- ## Rounds (full interview)
20
+ ## Full-interview rounds
17
21
 
18
- Work the decision tree from the top:
22
+ 1. Investigate first; list open decisions; keep only the frontier.
23
+ 2. Ask the round (rule 3).
24
+ 3. Record answers as one-line facts — `Settled: dark theme only`. Skipped questions settle to their stated default.
25
+ 4. Repeat until a round yields no new frontier questions.
26
+ 5. Close: print the decision list plus Given/When/Then acceptance examples for each load-bearing behaviour, then build (or hand back if the user only wanted the plan).
19
27
 
20
- 1. **Ask only the frontier.** A question belongs on the frontier when every decision it depends on is already settled. Asking "which database?" before "does this need persistence?" wastes a round; so does re-asking anything an earlier round settled.
21
- 2. **Batch the round.** Every frontier question in one message, numbered, each with `→ recommended: X` and a one-line reason. A recommendation is cheap for the user to confirm and expensive for them to derive.
22
- 3. **Never ask for facts.** Facts are yours: read the code, run the command, check the docs, delegate the wide search. If investigation can answer it, it is not a question — it is homework. Only decisions — taste, product calls, tradeoffs with real stakes — reach the user.
23
- 4. **Record what settled.** After each round, restate the settled decisions as one-line facts ("Settled: dark theme only") before asking the next round, so the record is unambiguous.
24
- 5. **Stop at empty frontier.** When a round produces no new questions, the session is done: restate the full decision list, then start building. The goal is settled decisions, not exhaustive documentation.
28
+ Hand-offs (shared-language skill): newly settled domain terms → glossary; a decision that is hard to reverse, surprising, and a real tradeoff → offer an ADR.
25
29
 
26
- ## Anti-patterns
30
+ ## Do not
27
31
 
28
- - **Interrogation drip** — one question per reply across ten replies. Every reply costs the user a context switch; batching is the fix.
29
- - **Homework outsourcing** — asking what you could have read. Burns trust and returns worse answers than the source.
30
- - **Speculative depth** — questions about futures nobody has committed to. Ask when the decision is load-bearing for work you are about to do.
32
+ - Drip one question per reply — batch the round.
33
+ - Ask about futures nobody has committed to; ask only what is load-bearing for work about to start.
34
+
35
+ ## Scaling: one agent or several
36
+
37
+ Main thread only. Fact-finding spanning several packages may go to `owl` (repo) or `researcher` (web) children in one `spawn_agent` call; the interview itself is never delegated.
38
+
39
+ Sources (accessed 3 October 2026): Given-When-Then — https://martinfowler.com/bliki/GivenWhenThen.html
@@ -1,30 +1,88 @@
1
1
  ---
2
2
  name: code-review
3
- description: Use when the user asks to review written work — a diff, PR, branch, or "review this before commit/merge" — checking both what was asked for and how well it is built. Do NOT use while still mid-build, when the user only wants a diff summary, or for security review — that is the bulletproof skill's lane.
3
+ description: Use when the user asks to review written work — a diff, PR, branch, commit range, or agent-generated change — or says "review before merge", "is this really done?", "check what the agent did". Checks spec compliance and build quality separately, verifies PR/"done" claims against the diff, and reports file:line findings by severity. Do NOT use while still mid-build (finish and run checks instead), for a plain diff summary, or for security review (bulletproof), performance (lean), migrations/data loss (durable), or UI review (evidence-led-ui).
4
4
  ---
5
5
 
6
6
  # Code Review
7
7
 
8
- Two independent questions, kept separate because each contaminates the other:
8
+ **Route first:**
9
9
 
10
- 1. **Spec axis** — does this change do what was asked, completely, with nothing asked that is missing?
11
- 2. **Standards axis** — does it meet how this repo builds: correctness, error handling, tests at real seams, naming, dead code, scope discipline?
10
+ | Situation | Mode | Next |
11
+ |---|---|---|
12
+ | Small diff: one concern, ≲ ~400 changed lines, few files | **Single pass** | Method below; both axes yourself, spec first, then standards |
13
+ | Large diff, multi-package PR, or many unrelated concerns | **Fan-out** | Build the ledger, then `## Scaling` |
14
+ | Diff written by an agent (including you, this session) | Either mode **+ agent checks** | Also run `references/agent-diffs.md` |
15
+ | User asks "is it secure?" | Defer | Load the `bulletproof` skill — do not audit security here |
16
+
17
+ Two questions, kept separate because each contaminates the other:
18
+
19
+ 1. **Spec axis** — does the change do what was asked, completely, and nothing unasked?
20
+ 2. **Standards axis** — does it meet how this repo builds: design, correctness, error handling, tests, naming, dead code, scope?
12
21
 
13
22
  ## Method
14
23
 
15
- 1. **Pre-flight.** Resolve the exact ref and range and confirm the diff is non-empty BEFORE any review work — a bad ref must fail here, not two passes deep. Read the request or task that motivated the change first: spec-conformance cannot be judged without the spec.
16
- 2. **Split the passes.** On anything larger than a small diff, run the two axes as separate subagents — the spec brief pastes in the original request, the standards brief pastes in the repo's conventions. Separate contexts keep one axis from anchoring the other. Small diffs: run both passes yourself, in that order, never interleaved.
17
- 3. **Findings format.** Per finding: `file:line`, one-line problem, why it matters, the concrete fix. Label each **[spec]** or **[standards]**, and never merge the lists into one ranking — the axes are not comparable, and merging re-ranks by noise.
18
- 4. **Report, then fix what is selected.** Report first; fix what the user picks. Never auto-apply fixes mid-review.
24
+ 1. **Pre-flight.** Resolve the exact ref/range (`git diff --stat <base>...<head>`) and confirm it is non-empty BEFORE review work — a bad ref must fail here. Read the motivating request/issue/task first; spec compliance cannot be judged without the spec. No spec available → say so and review standards only.
25
+ 2. **Claims ledger.** List every claim from the PR description, commit messages, or the agent's "done" report ("adds X", "fixes Y", "tests pass", "no behaviour change"). Each row ends `verified (file:line or command output)` / `contradicted` / `unverifiable`. A claim with no matching hunk is a **[spec]** finding.
26
+ 3. **Run what is cheap.** Typecheck, lint, affected tests — the commands CI uses. Record the actual result. Never write "tests pass" from reading.
27
+ 4. **Read the whole diff, then the context.** Every changed line, plus callers of changed signatures (`code_nav references`). Start with design: does the change belong here, at this layer, now? (Google eng-practices ranks design first.)
28
+ 5. **Spec pass, then standards pass.** Never interleave. Standards baseline below; agent failure modes in `references/agent-diffs.md`; test quality in `references/tests.md`.
29
+ 6. **Report, then fix what the user selects.** Never auto-apply fixes mid-review.
19
30
 
20
- Security findings belong to the `bulletproof` skill — note them and defer; do not audit them here.
31
+ ## Findings format
21
32
 
22
- ## Standards baseline (when the repo defines none)
33
+ One line each, anchored and actionable:
34
+
35
+ `[severity] [spec|standards] path/file.ts:42 — problem. Why it matters. Fix: concrete change.`
23
36
 
24
- Missing error handling on I/O and external calls; load-bearing logic without tests, or tests asserting implementation details; dead code and commented-out blocks; names that hide intent; an abstraction used once; a new dependency where the standard library or an installed package would do; secrets in the diff; suppression of failing checks (skipped tests, `as any`, relaxed assertions).
37
+ | Severity | Means | Example |
38
+ |---|---|---|
39
+ | **blocking** | Must change before merge | Wrong behaviour, missing requirement, deleted/weakened test, swallowed error, hallucinated API or package |
40
+ | **non-blocking** | Should change; may follow up | Duplicate helper, weak assertion on a non-critical path |
41
+ | **nit** | Optional polish | Naming, local readability |
42
+ | **question** | You cannot tell; author must answer | Unclear intent, unverifiable claim |
43
+
44
+ Rules:
45
+ - Keep **[spec]** and **[standards]** as two separate lists — never merge the lists into one ranking; the axes are not comparable and merging re-ranks by noise. Order by severity within each list.
46
+ - Every finding cites a `file:line` you re-opened yourself. No line → it is a **question**, not a finding.
47
+ - Comment on the code, not the author; say why.
48
+ - Skip what tooling already enforces — if lint/CI catches it, it is noise.
49
+ - Security: one line, `[defer → bulletproof] file:line — what looked risky`. Do not rate or fix it here.
50
+
51
+ ## Standards baseline (when the repo defines none)
25
52
 
26
- Skip anything tooling already enforces — findings the linter or CI catches are noise.
53
+ - **Design:** abstraction with one caller; logic in the wrong layer; new dependency where stdlib or an installed package does it.
54
+ - **Correctness:** boundary/off-by-one, null/empty handling, ordering/concurrency changes, error paths.
55
+ - **Errors:** I/O and external calls without handling; broad `catch` that swallows or logs-and-continues.
56
+ - **Tests:** behaviour change without a test; tests that assert internals or nothing (`references/tests.md`).
57
+ - **Hygiene:** dead code, commented-out blocks, debug prints, names that hide intent, secrets in the diff.
58
+ - **Suppression:** skipped tests, `as any`/`@ts-ignore`/`noqa`/`eslint-disable`, relaxed thresholds, `continue-on-error` in CI.
59
+ - **Scope:** files or refactors outside the request; config, lockfile, or CI changes nobody asked for.
27
60
 
28
61
  ## Verdict
29
62
 
30
- End with the honest state: **works as asked** / **works with gaps** (name them) / **not ready** (blocking reasons). Never "looks good" without stating what was and was not verified.
63
+ - **works as asked** / **works with gaps** (name them) / **not ready** (blocking reasons).
64
+ - Claims ledger: N verified, N contradicted, N unverifiable.
65
+ - What you ran (commands + result) and **what was not checked**.
66
+
67
+ Never "looks good"/"LGTM" without stating what was and was not verified. Never certify a change safe or bug-free.
68
+
69
+ ## Scaling: one agent or several
70
+
71
+ | Situation | Do |
72
+ |---|---|
73
+ | Single-pass size (one concern, ≲ ~400 changed lines, few files) | Main thread only. Never spawn. |
74
+ | Larger, multi-package, or > ~15 files | Main thread builds a **coverage ledger**: rows = lens (spec, standards/correctness, tests, agent checks) × file group (package/directory). Each row ends `checked-with-findings` / `checked-clean` / `not-checked(reason)`. |
75
+ | Fan-out | ONE `spawn_agent` call, ≤ 6 read-only children per lens and/or file group: `owl` for reading; general-purpose child if it must run tests. |
76
+ | Security-sensitive hunks (auth, input, secrets, deps, CI) | Route to bulletproof's protocol: `auditor` child briefed with the bulletproof skill root, then `skeptic` on its findings. |
77
+ | High-stakes merge (release, > ~1,000 lines) | One fresh-context verifier child tries to disprove each blocking finding. |
78
+
79
+ **Child brief** (children see nothing else): absolute skill root `…/assets/skills/code-review`; reference file(s) to read; the exact `git diff <base>...<head> -- <paths>` to run; its ledger rows; the spec text verbatim (spec lens) or the repo conventions (standards lens); the output schema — findings in the format above plus explicit `checked` and `not checked` lists. Read-only, no edits.
80
+
81
+ **Merge:** a child that fails, times out, or omits a row → that row is `not checked`, never clean. Re-open every reported `file:line` before reporting it; drop what you cannot confirm. Dedupe across children. Keep spec and standards lists separate. Fixes afterwards are serialized in the main thread.
82
+
83
+ ## Sources (SNAPSHOT, accessed 3 October 2026)
84
+
85
+ - Google eng-practices — https://google.github.io/eng-practices/review/reviewer/looking-for.html, …/reviewer/comments.html, …/reviewer/standard.html, …/developer/small-cls.html
86
+ - Conventional Comments (labels, blocking/non-blocking) — https://conventionalcomments.org/
87
+ - Mutation-testing concept — https://stryker-mutator.io/docs/
88
+ - Package hallucination — Spracklen et al., USENIX Security 2025 (arXiv:2406.10279); 2026 re-evaluation, not peer-reviewed (arXiv:2605.17062), via https://socket.dev/blog/slopsquatting-targets-across-frontier-llms
@@ -0,0 +1,27 @@
1
+ # Agent-Generated Diffs: Failure Modes and Checks
2
+
3
+ Run on any diff an agent wrote (including your own). Each row: what to look for → how to check → default severity.
4
+
5
+ | Failure mode | Check | Severity |
6
+ |---|---|---|
7
+ | **Hallucinated package** | Every new dependency in a manifest: confirm it exists in the registry (`npm view <pkg>`, `pip index versions <pkg>`), is the intended project (not a look-alike name), and is actually imported. Install scripts → defer to bulletproof. | blocking |
8
+ | **Hallucinated API / symbol** | Every newly called function, method, option, or flag: `code_nav definition`, or read the installed source in `node_modules`/site-packages for the pinned version. Typecheck passing is evidence only for typed code. | blocking |
9
+ | **Weakened or deleted tests** | `git diff --stat -- '*test*' '*spec*'`. For each test hunk: removed assertions, loosened matchers (`toEqual`→`toBeDefined`, exact→`toContain`), widened tolerances, added `.skip`/`.only`/`xit`/`@pytest.mark.skip`, snapshots re-recorded without reason. | blocking unless justified in the PR |
10
+ | **Mocked-away assertions** | Test mocks the very unit under test, or asserts only that a mock was called with whatever it was given. | blocking on critical paths |
11
+ | **Swallowed errors** | New `try/catch` / `except Exception` that returns a default, logs and continues, or wraps a whole function. Compare with the code's previous error contract. | blocking |
12
+ | **Dead code / duplicate helpers** | New function that duplicates an existing helper (`grep` the name's verbs, `code_search` the behaviour); unused exports, params, or branches left behind. | non-blocking |
13
+ | **Scope creep / unrequested refactor** | Map every changed file to a requirement. Files with no requirement → list them. Renames/reformatting mixed with behaviour change → ask for a split. | non-blocking; blocking if it changes behaviour |
14
+ | **Fake "done"** | Claims ledger (SKILL.md step 2): "tests pass" with no run; "handles X" with no hunk; TODO/placeholder/`throw new Error("not implemented")` left in. | blocking |
15
+ | **Config / lockfile drift** | Lockfile changes without a manifest change (or vice versa); tsconfig/eslint/biome rules relaxed; CI steps removed or `continue-on-error` added; version bumps nobody asked for. | blocking for relaxed checks; question otherwise |
16
+ | **Suppressions** | New `any`, `!`, `@ts-ignore`, `eslint-disable`, `# type: ignore`, `noqa`. Each needs a reason at the line. | non-blocking (blocking if hiding a failing check) |
17
+ | **Generated / vendored edits** | Hand edits in `dist/`, generated, or vendored files. | blocking |
18
+
19
+ ## Quick commands
20
+
21
+ ```bash
22
+ git diff --stat <base>...<head> # scope
23
+ git diff <base>...<head> -- '*.lock' '*lock.json' '*.toml' 'package.json' # dependency drift
24
+ git diff <base>...<head> | grep -nE '^\+.*(\.skip|\.only|xit\(|@ts-ignore|eslint-disable|as any|catch *\()'
25
+ ```
26
+
27
+ Grep hits are leads, not findings — open each at file:line before reporting.
@@ -0,0 +1,19 @@
1
+ # Reviewing Tests
2
+
3
+ Coverage says a line ran, not that a test would notice it breaking. Review tests with a mutation-testing mindset: "if I broke this line, which test fails?"
4
+
5
+ ## Checks
6
+
7
+ 1. **Each behaviour change has a test that fails without it.** Mentally revert the production hunk and trace which assertion would fail. None → the test does not cover the change. (Do not stash or reset the author's work to check; if a run is needed, ask first.)
8
+ 2. **Mutation spot-check on critical lines** (money, auth, data writes, parsing boundaries): flip a comparison, drop a condition, return early. If a test framework exists (e.g. StrykerJS, mutmut, PIT), suggest running it on the changed files only; do not install it unasked.
9
+ 3. **Assertion quality.** Flag:
10
+ - no assertion, or only `toBeDefined`/`toBeTruthy`/`not.toThrow` where a value is knowable;
11
+ - asserting a mock was called, with no check of the outcome;
12
+ - snapshot of a huge object where one field matters;
13
+ - assertions on private internals that block refactoring.
14
+ 4. **Real seams.** Mocks only at external boundaries (network, clock, filesystem, randomness). Mocking the unit under test, or the module next to it, proves nothing.
15
+ 5. **Edge cases present:** empty, null/undefined, boundary values, error path, concurrency/ordering when relevant.
16
+ 6. **Determinism.** No real time, randomness, network, or test-order dependence. Flaky = blocking.
17
+ 7. **Test diffs in a "no behaviour change" PR** need a stated reason; otherwise treat as weakened tests (see `agent-diffs.md` table).
18
+
19
+ Report test findings under **[standards]** unless a required test the spec asked for is missing — that is **[spec]**.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: compliance-guard
3
- description: Use when shipping something real users reach and the work carries legal exposure — pre-launch or "is this safe to ship" reviews; personal data, tracking pixels, cookies and consent; payments, subscriptions, auto-renewal; user uploads or UGC; email/SMS; AI or chatbot features; minors' data; biometrics; scraping; accessibility of public pages; or drafting privacy policies, terms, and disclosures. Also use when a feature may be licensed or illegal: health, finance, legal advice, money movement, crypto, gambling, adult content, background checks, automated hiring/lending/housing decisions. Do NOT use for local-only scripts, throwaway prototypes with no real users or data, or changes with no data, money, users, or public surface.
3
+ description: Use when shipping something real users reach and the work carries legal exposure — pre-launch or "is this safe to ship" reviews; mid-build features adding personal data, pixels/cookies, payments or auto-renewal, uploads/UGC, email/SMS/push, AI chatbots, minors, biometrics, scraping; drafting privacy policies, terms, disclosures. Also when a feature may need a licence or be illegal: health, finance, money movement, crypto, gambling, adult content, background checks, automated hiring/lending/housing decisions. Do NOT use for local-only scripts, prototypes with no real users or data, or changes touching no data, money, users or public surface; security hardening is bulletproof's lane.
4
4
  license: Apache-2.0. Content is engineering guidance, not legal advice. See references/provenance.md.
5
5
  compatibility: Static review works offline from the bundled references. Legal status changes constantly; date-sensitive claims must be re-verified with web access before being stated as current. Never certifies compliance.
6
6
  ---
@@ -13,9 +13,9 @@ Catch the legal, privacy, and regulatory exposure in a shipped product before a
13
13
 
14
14
  1. **Exposure drives obligations, not stack.** What the product *does*, *who can reach it*, and *whose data it touches* decide what applies. A CLI that never leaves the laptop owes almost nothing. A one-page site with a contact form and an ad pixel owes a surprising amount.
15
15
  2. **Never certify.** Do not write or say "compliant", "GDPR compliant", "ADA compliant", "fully legal", or "you're covered". Produce a risk register, implemented controls, residual risk, and an explicit *get a lawyer for this* list. This is engineering guidance, not legal advice, and must be labelled as such in every report.
16
- 3. **Date-check before asserting.** The references are a snapshot dated **11 August 2026**. Effective dates, thresholds, injunctions, and penalty amounts move. Before stating a date, a threshold, or "this is in force", re-verify with web access if available; if unavailable, say the claim is from a dated snapshot and needs confirmation. Never invent a citation, statute section, or deadline.
16
+ 3. **Date-check before asserting.** The references are a snapshot dated **11 August 2026**, spot re-verified **3 October 2026** (`references/provenance.md` lists what was and was not re-checked). Effective dates, thresholds, injunctions, and penalty amounts move. Before stating a date, a threshold, or "this is in force", re-verify with web access if available; if unavailable, say the claim is from a dated snapshot and needs confirmation. Never invent a citation, statute section, or deadline.
17
17
  4. **Say it plainly when it is illegal.** If the requested build is unlawful, licensed, or criminal as described, state that clearly and early — before writing code, not after. Name the specific regime, the concrete red line, the safe subset that *can* be built, and what authorization would change the answer. Do not soften it into a vague caution, and do not silently build it.
18
- 5. **Fix, do not just flag.** Anything code can fix, fix: consent gating, security P0s, opt-out plumbing, deletion propagation, disclosure strings, accessibility defects. Draft documents as clearly-marked templates with `[PLACEHOLDER]` fields. Never invent the user's legal facts — entity name, registered address, DPO, retention periods, or vendor list must come from the user or the repo.
18
+ 5. **Fix, do not just flag.** Anything code can fix, fix: consent gating, security P0s (via bulletproof's inline gate), opt-out plumbing, deletion propagation, disclosure strings, accessibility defects. Draft documents as clearly-marked templates with `[PLACEHOLDER]` fields. Never invent the user's legal facts — entity name, registered address, DPO, retention periods, or vendor list must come from the user or the repo.
19
19
  6. **Proportionality.** A weekend prototype with no users does not need 60 findings. Gate on *launch-blocking* first, rank by probability × severity, and keep the tail as a backlog. Overwhelming a solo dev produces zero fixes.
20
20
  7. **Jurisdictions are a matrix, not a country.** "Where the company is" rarely limits exposure; "who can reach the app" usually sets it. A US-only startup with EU visitors and an unblocked signup form is in scope for EU law.
21
21
 
@@ -25,7 +25,7 @@ Catch the legal, privacy, and regulatory exposure in a shipped product before a
25
25
 
26
26
  This mode matters most, because the users who need this skill will never ask for it. They ask for a signup page, a Stripe checkout, a contact form, an image upload. **Build the control into the feature as you write it** rather than waiting to be asked — consent-gate the pixel you were told to add, put the unsubscribe link in the email template, enforce authorization at the data layer. Mention it in one line and move on. Do not stop the build to deliver a lecture, and do not silently ship the unsafe version and flag it later.
27
27
 
28
- **Full review** — triggered by pre-launch, "is this safe", audit, or first use of this skill on a project. Run the whole workflow below and write the register.
28
+ **Full review** — triggered by pre-launch, "is this safe", or an audit. A first use on a project mid-build stays an inline gate. Run the whole workflow below and write the register.
29
29
 
30
30
  ## Workflow
31
31
 
@@ -214,6 +214,21 @@ The register is the artifact; the message to the user is what actually gets acte
214
214
 
215
215
  **Do not moralise and do not catastrophise.** State the base rate. "Small apps get sued over this regularly" and "this almost never gets enforced against someone your size, but it is cheap to fix" are both useful; "you could be fined €20 million" is not, because they will stop reading.
216
216
 
217
+ ## Scaling: one agent or several
218
+
219
+ | Situation | Do |
220
+ |---|---|
221
+ | Inline gate, small edit, single feature | Main thread only. Never spawn for an inline gate. |
222
+ | Full review, one deployable, one primary jurisdiction, every ledger row readable by you | Stay single-threaded. |
223
+ | Several deployables/surfaces (web + mobile + admin + marketing site) **and** several jurisdictions or data flows | Build the coverage ledger first (rows = ledger items × surfaces × jurisdictions), then fan out. |
224
+
225
+ Fan-out rules:
226
+ - One `spawn_agent` call, ≤ 6 read-only children, each owning a disjoint slice — **by product surface** (paths) or **by jurisdiction** (US / EU / UK). Brief each with: absolute skill root, the exact reference files to read, its paths, the ledger rows it owns, the RUNTIME/CODE/DEDUCED labels, `[V]/[S]/[U]` markers, and the output schema (finding, file:line, label, severity, fix, plus explicit `checked` and `not checked` lists). Children never see this conversation.
227
+ - Dated legal claims the report will rely on → one `researcher` child per jurisdiction to re-verify against primary sources (statute site, official journal, regulator, court), returning date, status, and URL.
228
+ - **Merge:** a child that fails, times out, or omits a row → that row is `not checked`, never clean. Re-open each reported file:line yourself before reporting it.
229
+ - **Independent verification, always, for a full review:** a fresh-context child gets the draft findings and tries to disprove each one (wrong jurisdiction, ban vs duty, stale date, wrong label). Downgrade or drop what it refutes.
230
+ - Fixes stay serialized in the main thread (or `bee` children on strictly disjoint files); run checks once after merging.
231
+
217
232
  ## Hard stops
218
233
 
219
234
  Do not build these, regardless of framing. State the reason and offer the lawful subset from `references/sector-gates.md`:
@@ -232,7 +247,7 @@ Do not build these, regardless of framing. State the reason and offer the lawful
232
247
  - Never present something you read as something you ran. Every finding and every fix is **RUNTIME**, **CODE**, or **DEDUCED**, and the report says which. "I could not verify this" is a legitimate and useful output; a fabricated confirmation is not.
233
248
  - Report what you did **not** check. A review that silently skips the mobile app, the admin panel, or the marketing site reads as full coverage and is more dangerous than no review.
234
249
  - Never fabricate a statute, section number, case, effective date, or penalty. If unsure, say "verify this" and mark confidence.
235
- - Distinguish **verified**, **snapshot (11 Aug 2026, re-verify)**, and **uncertain** in the report. The references carry these markers — preserve them; do not launder a flagged-uncertain item into a confident statement.
250
+ - Distinguish **verified**, **snapshot (11 Aug 2026 / 3 Oct 2026, re-verify)**, and **uncertain** in the report. The references carry these markers — preserve them; do not launder a flagged-uncertain item into a confident statement.
236
251
  - Do not use fear as a lever. Give the base rate and the fix, not doom.
237
252
  - When the user says a jurisdiction does not apply to them, record it as their stated assumption rather than silently accepting or arguing.
238
253
 
@@ -1,6 +1,6 @@
1
1
  # Artifacts
2
2
 
3
- What each required document or flow must actually contain, as a skeleton to generate against. Snapshot **11 Aug 2026**.
3
+ What each required document or flow must actually contain, as a skeleton to generate against. Snapshot **11 Aug 2026**, spot re-verified 3 Oct 2026 (see `provenance.md`).
4
4
 
5
5
  **Rule for every document here:** generate it with `[PLACEHOLDER]` fields for facts only the user has, never invent entity names, addresses, retention periods, or vendor lists, and mark the output as a template that needs review. State plainly which artifacts a competent developer can safely template and which genuinely need a lawyer.
6
6
 
@@ -1,6 +1,6 @@
1
1
  # EU / EEA / UK
2
2
 
3
- Snapshot **11 Aug 2026**. Markers: **[V]** verified against a primary/first-tier source · **[S]** snapshot, re-verify · **[U]** contested or in flux.
3
+ Snapshot **11 Aug 2026**, spot re-verified 3 Oct 2026 (see `provenance.md`). Markers: **[V]** verified against a primary/first-tier source · **[S]** snapshot, re-verify · **[U]** contested or in flux.
4
4
 
5
5
  **Territorial reality:** GDPR applies to anyone offering goods or services to, or *monitoring the behaviour of*, people in the EU/EEA — regardless of where the developer sits **[V]**. Monitoring includes analytics cookies, session replay, ad pixels, and behavioural profiling. A US solo dev with an open signup form and Google Analytics is in scope. Blocking EU traffic is a legitimate engineering answer and should be offered as an option.
6
6
 
@@ -38,7 +38,7 @@ Snapshot **11 Aug 2026**. Markers: **[V]** verified against a primary/first-tier
38
38
 
39
39
  ## 2. ePrivacy — cookies and device storage
40
40
 
41
- Consent is required **before** any non-essential storage or access on the user's device, under the ePrivacy Directive as transposed by each Member State — 27 variants, and harmonisation is not coming **[V]**. Scope is technology-neutral and expressly reaches pixels, local storage, and similar techniques **[V]**. This applies **whether or not** the data is personal, which is why "we only use anonymous analytics" is not an answer.
41
+ Consent is required **before** any non-essential storage or access on the user's device, under the ePrivacy Directive as transposed by each Member State — 27 variants **[V]**. The Digital Omnibus proposal to move cookie rules into the GDPR (Art 88a/88b: one-click refuse, no re-asking for six months, browser signals) is **not law** — as of September 2026 it was still in first reading with no Council negotiating mandate and no trilogue **[S]**. Do not build to it; one-click reject is already the defensible design anyway. Scope is technology-neutral and expressly reaches pixels, local storage, and similar techniques **[V]**. This applies **whether or not** the data is personal, which is why "we only use anonymous analytics" is not an answer.
42
42
 
43
43
  **Banner technical spec** (this is the implementable contract):
44
44
  - No third-party script, pixel, or network request before an affirmative choice.
@@ -60,19 +60,19 @@ Consent is required **before** any non-essential storage or access on the user's
60
60
 
61
61
  ---
62
62
 
63
- ## 3. EU AI Act — verified timeline
63
+ ## 3. EU AI Act — timeline
64
64
 
65
- The Digital Omnibus on AI (Reg (EU) 2026/1744) was published on 24 July 2026 and entered into force on 27 July 2026 **[V]**. **Any guidance dated before mid-2026 saying high-risk obligations apply from 2 August 2026 is now wrong.**
65
+ The Digital Omnibus on AI (Reg (EU) 2026/1744 of 8 July 2026) was published in the OJ on 24 July 2026 and entered into force on the third day after publication **[V]** (EUR-Lex OJ text opened 3 Oct 2026). Dates below marked **[V]** were read in that text; 2 Feb 2027 was not located there **[U]**. **Any guidance dated before mid-2026 saying high-risk obligations apply from 2 August 2026 is now wrong.**
66
66
 
67
67
  | Date | Status |
68
68
  |---|---|
69
- | 2 Feb 2025 | Prohibited practices and AI literacy — **in force** |
70
- | 2 Aug 2025 | GPAI models, governance, penalties — **in force** |
71
- | **2 Aug 2026** | **Article 50 transparency obligations — in force now** (Art 50(2) not applying to systems already on the market at that date) |
72
- | 2 Dec 2026 | Art 50(2) marking for legacy systems; new prohibited practices added (AI-generated NCII and CSAM) |
73
- | 2 Feb 2027 | Watermark-detection interoperability deadline for providers |
74
- | **2 Dec 2027** | High-risk obligations for standalone Annex III systems — **deferred** |
75
- | 2 Aug 2028 | High-risk obligations for embedded Annex I systems |
69
+ | 2 Feb 2025 | Prohibited practices and AI literacy — **in force** **[V]** |
70
+ | 2 Aug 2025 | GPAI models, governance, penalties — **in force** **[S]** |
71
+ | **2 Aug 2026** | **Article 50 transparency obligations — in force now** (general application date) **[V]** |
72
+ | 2 Dec 2026 | Art 50(2) marking for generative systems placed on the market before 2 Aug 2026; new Art 5(1)(ba)/(bb) prohibitions incl. non-consensual sexual deepfakes of identifiable people **[V]** |
73
+ | 2 Feb 2027 | Watermark-detection interoperability deadline for providers **[U]** |
74
+ | **2 Dec 2027** | High-risk obligations for standalone Annex III systems — **deferred** **[V]** |
75
+ | 2 Aug 2028 | High-risk obligations for embedded Annex I systems **[V]** |
76
76
 
77
77
  **Provider vs deployer is the consequential classification.** Providers develop an AI system, or have it developed, and place it on the market or into service **under their own name or trademark**, regardless of establishment **[V]**. Wrapping a foundation model in your own product and shipping it under your brand generally makes you a provider of that AI system — not merely a deployer.
78
78
 
@@ -114,10 +114,10 @@ Penalties for Art 50 breaches reach €15M or 3% of global turnover; prohibited
114
114
  ## 4. Other EU acts
115
115
 
116
116
  - **Digital Services Act** — triggered by *hosting information provided by a recipient*: user uploads, comments, profiles, public pastes, shared docs. Single-tenant B2B SaaS with no third-party-visible content is generally out **[V]**. All hosting providers regardless of size owe a point of contact for authorities and users, a notice-and-action mechanism, statements of reasons for removals, and terms describing moderation. Micro and small enterprises are exempt from several heavier duties but **not** from the basics.
117
- - **European Accessibility Act** — applies to **service categories**, not all software: e-commerce (any consumer-facing online sale), consumer banking, e-books, electronic communications, transport ticketing, and access to audiovisual media **[V]**. In force since 28 June 2025 for new products and services. A B2B-only SaaS is out of scope; a B2C app with a checkout is in. Technical standard is EN 301 549 (WCAG 2.1 AA today; a WCAG 2.2-aligned version is expected **[U]** — build to 2.2 AA now, it is backwards-compatible). Microenterprise exemptions apply to services but the detail varies by transposition **[U]**.
117
+ - **European Accessibility Act (Directive (EU) 2019/882)** — applies to **service categories**, not all software: e-commerce (any consumer-facing online sale), consumer banking, e-books, electronic communications, transport ticketing, and access to audiovisual media **[V]**. Applies since 28 June 2025 to products placed on the market and services provided to consumers after that date (Art 31(2)) **[V]**. Art 32 transitional rules: service contracts agreed before 28 June 2025 may run unaltered until expiry but no later than **28 June 2030**; products already used to deliver services may continue until 28 June 2030; self-service terminals up to 20 years **[S]** (Article text as quoted by secondary sources; EUR-Lex not opened). None of this exempts a new or changed web checkout. A B2B-only SaaS is out of scope; a B2C app with a checkout is in. Technical standard is EN 301 549 (WCAG 2.1 AA today; a WCAG 2.2-aligned revision is in progress **[U]** — build to 2.2 AA now, it is backwards-compatible). Microenterprises (<10 staff and ≤€2M turnover) are exempt for services, but the detail varies by transposition **[S]**.
118
118
  - **Cyber Resilience Act** — applies to *manufacturers* of products with digital elements placed on the EU market: downloadable or installable software, desktop and mobile apps, browser extensions, firmware, monetised libraries **[V]**. Main obligations from 11 December 2027, but **reporting obligations from 11 September 2026** — actively exploited vulnerabilities and severe incidents must be reported on a short clock. Pure SaaS is generally outside, but the SaaS boundary is exactly where small products get caught unexpectedly; re-read the Commission's practical guidance **[U]**.
119
119
  - **NIS2** — sector plus size; cloud, data-centre, managed-service and managed-security providers are in scope, but the size cap generally means ≥50 staff or >€10M turnover **[S]**. A small SaaS is normally out unless designated.
120
- - **Data Act** — applies to providers of data processing services (expressly including SaaS, PaaS, IaaS) with EU customers, with no carve-out for small providers **[V]**. Practical duties: contractual switching and egress terms, no unreasonable exit barriers, and data-porting support.
120
+ - **Data Act (Regulation (EU) 2023/2854)** — applicable since **12 September 2025** (Art 50) **[S]**. Two developer-relevant parts: (1) providers of data processing services (expressly including SaaS, PaaS, IaaS) with EU customers, with no carve-out for small providers, owe contractual switching and egress terms, no unreasonable exit barriers, and data-porting support; switching charges must be abolished from **12 January 2027** **[S]**. (2) Makers of **connected products** (IoT devices, wearables, connected cars) and their companion apps owe user access to product and related-service data; from **12 September 2026** products placed on the market must be designed so users can access that data directly where relevant and technically feasible (Art 3(1)) **[S]**. Chapter IV unfair-terms rules apply to B2B data contracts concluded after 12 Sept 2025 **[S]**.
121
121
  - **DORA** — only if you are a financial entity or a contracted ICT provider to one; for a small dev it arrives as customer contract terms **[S]**.
122
122
  - **PSD2 SCA** — use a PSP with 3-D Secure rather than building card flows; exemptions belong to the PSP **[S]**.
123
123
  - **MiCA** — issuing a token or providing crypto-asset services to EU users requires authorisation; merely accepting crypto payment through a licensed processor generally does not **[S]**.
@@ -128,8 +128,8 @@ Penalties for Art 50 breaches reach €15M or 3% of global turnover; prohibited
128
128
 
129
129
  ## 5. UK specifics
130
130
 
131
- - **Data (Use and Access) Act 2025** — principal data-protection provisions in force from 5 February 2026 **[V]**. Code-relevant changes: the new DSAR clock (Art 12A), recognised legitimate interests, a permission-plus-safeguards model for automated decision-making replacing the old prohibition, and the PECR analytics/functionality cookie exemption.
132
- - **Online Safety Act** — triggered by **user-to-user** services (anywhere users can encounter content uploaded by others — comments, DMs, forums, shared galleries, multiplayer chat), search services, and pornography publishers, with UK links. **There is no small-service exemption from the core duties**, and the regulator runs a dedicated "small but risky" supervision function **[V]**. Duties include illegal-content and children's-access risk assessments, proportionate safety measures, reporting and complaints mechanisms, and highly effective age assurance where required. Treat any UK-reachable UGC product as in scope and produce the risk assessments — their absence is itself the enforceable failure.
131
+ - **Data (Use and Access) Act 2025** — Royal Assent 19 June 2025; commenced in stages. The principal data-protection provisions came into force on **5 February 2026** under SI 2026/82 (legislation.gov.uk): ss. 70, 76, 80, 112 and Schedules 4, 6, 11, 12 **[V]**. Code-relevant changes: the new DSAR clock (Art 12A UK GDPR), recognised legitimate interests (no balancing test), Arts 22A–22D replacing the automated-decision prohibition with permission plus safeguards (information, representations, human intervention, contest) for non-special-category data, and the PECR analytics/functionality cookie exemption. The complaints-handling duty (new s. 164A DPA 2018) followed on **19 June 2026** **[V]**. Remaining stages (ICO restructuring into the Information Commission) are institutional and do not change code **[S]**.
132
+ - **Online Safety Act** — triggered by **user-to-user** services (anywhere users can encounter content uploaded by others — comments, DMs, forums, shared galleries, multiplayer chat), search services, and pornography publishers, with UK links. **There is no small-service exemption from the core duties**, and the regulator runs a dedicated "small but risky" supervision function **[V]**. Duties include illegal-content and children's-access risk assessments, proportionate safety measures, reporting and complaints mechanisms, and highly effective age assurance where required. Ofcom has been fining since late 2025 — mostly for missing highly effective age assurance on pornographic services, and in one case for a missing illegal-content risk assessment and deficient terms — with penalties up to £18M or 10% of qualifying worldwide revenue **[S]**. The register of categorised services (Category 1, 2A, 2B; extra duties) was published in July 2026 **[S]**; small services are not categorised, but the core duties still apply. Treat any UK-reachable UGC product as in scope and produce the risk assessments — their absence is itself the enforceable failure.
133
133
  - **Children's Code** — applies to services *likely to be accessed* by under-18s, a much lower bar than "aimed at children": high-privacy defaults, geolocation off by default, no nudges toward weaker privacy, a DPIA covering children, minimised profiling **[V]**.
134
134
  - **Accessibility** — no private-sector EAA equivalent; exposure runs through the Equality Act duty to make reasonable adjustments, with WCAG 2.1 AA as the de facto benchmark **[S]**.
135
135
 
@@ -148,4 +148,4 @@ Penalties for Art 50 breaches reach €15M or 3% of global turnover; prohibited
148
148
 
149
149
  ## 7. Verify before relying
150
150
 
151
- AI Act fine tiers post-Omnibus · the EAA transitional date for pre-2025 service contracts (sources conflict between 2027 and 2030) · EN 301 549 version status · the GDPR/ePrivacy half of the Digital Omnibus (**still in negotiation — not law**) · consent-or-pay scope broadening · CRA SaaS boundary · NIS2 small-provider designation practice · UK adequacy status.
151
+ AI Act fine tiers post-Omnibus · EN 301 549 version status · the GDPR/ePrivacy half of the Digital Omnibus (**not law**: as of Sept 2026 still in first reading, no Council mandate, no trilogue) · Ofcom's categorised-services duties consultation outcome · consent-or-pay scope broadening · CRA SaaS boundary · NIS2 small-provider designation practice · UK adequacy status.