@proflandrigan/shards 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (397) hide show
  1. package/README.md +475 -0
  2. package/package.json +37 -0
  3. package/src/agents/academic.md +276 -0
  4. package/src/agents/ai-engineer.md +377 -0
  5. package/src/agents/analytics-engineer.md +364 -0
  6. package/src/agents/applied-ml-scientist.md +410 -0
  7. package/src/agents/backend-engineer.md +255 -0
  8. package/src/agents/bi-engineer.md +333 -0
  9. package/src/agents/data-analyst.md +343 -0
  10. package/src/agents/data-engineer.md +260 -0
  11. package/src/agents/data-modeller.md +386 -0
  12. package/src/agents/data-scientist.md +366 -0
  13. package/src/agents/deep-learning-engineer.md +389 -0
  14. package/src/agents/ml-engineer.md +424 -0
  15. package/src/agents/mlops-engineer.md +339 -0
  16. package/src/agents/researcher.md +187 -0
  17. package/src/agents/specific_instructions/academic/critical_review.md +263 -0
  18. package/src/agents/specific_instructions/academic/report.md +113 -0
  19. package/src/agents/specific_instructions/ai_engineer/advise.md +162 -0
  20. package/src/agents/specific_instructions/ai_engineer/bi_engineer_handoff.md +86 -0
  21. package/src/agents/specific_instructions/ai_engineer/experiment.md +471 -0
  22. package/src/agents/specific_instructions/ai_engineer/experiment_ui_mode.md +44 -0
  23. package/src/agents/specific_instructions/ai_engineer/phases/index.md +45 -0
  24. package/src/agents/specific_instructions/ai_engineer/phases/phase-1.md +55 -0
  25. package/src/agents/specific_instructions/ai_engineer/phases/phase-2.md +86 -0
  26. package/src/agents/specific_instructions/ai_engineer/phases/phase-3.md +96 -0
  27. package/src/agents/specific_instructions/ai_engineer/phases/phase-4.md +138 -0
  28. package/src/agents/specific_instructions/ai_engineer/phases/phase-5.md +157 -0
  29. package/src/agents/specific_instructions/ai_engineer/phases/phase-6.md +196 -0
  30. package/src/agents/specific_instructions/ai_engineer/phases/phase-7.md +313 -0
  31. package/src/agents/specific_instructions/ai_engineer/phases.md +1011 -0
  32. package/src/agents/specific_instructions/ai_engineer/prompt_lab.md +161 -0
  33. package/src/agents/specific_instructions/ai_engineer/prompt_lab_ui_mode.md +28 -0
  34. package/src/agents/specific_instructions/ai_engineer/research.md +393 -0
  35. package/src/agents/specific_instructions/ai_engineer/research_ui_mode.md +66 -0
  36. package/src/agents/specific_instructions/ai_engineer/review.md +159 -0
  37. package/src/agents/specific_instructions/ai_engineer/validation_checklist.md +182 -0
  38. package/src/agents/specific_instructions/analytics_engineer/advise.md +155 -0
  39. package/src/agents/specific_instructions/analytics_engineer/bi_engineer_handoff.md +91 -0
  40. package/src/agents/specific_instructions/analytics_engineer/data_analyst_handoff.md +84 -0
  41. package/src/agents/specific_instructions/analytics_engineer/deep_phases.md +818 -0
  42. package/src/agents/specific_instructions/analytics_engineer/phases_deep/index.md +24 -0
  43. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-1.md +77 -0
  44. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-2.md +106 -0
  45. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-3.md +93 -0
  46. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-4.md +79 -0
  47. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-5.md +61 -0
  48. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-6.md +45 -0
  49. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-7.md +235 -0
  50. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-8.md +221 -0
  51. package/src/agents/specific_instructions/analytics_engineer/phases_quick/index.md +19 -0
  52. package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-1.md +47 -0
  53. package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-2.md +78 -0
  54. package/src/agents/specific_instructions/analytics_engineer/quick_phases.md +112 -0
  55. package/src/agents/specific_instructions/analytics_engineer/review.md +167 -0
  56. package/src/agents/specific_instructions/analytics_engineer/service_mode.md +369 -0
  57. package/src/agents/specific_instructions/analytics_engineer/ui_mode.md +45 -0
  58. package/src/agents/specific_instructions/analytics_engineer/update.md +162 -0
  59. package/src/agents/specific_instructions/analytics_engineer/validation_checklist.md +121 -0
  60. package/src/agents/specific_instructions/applied_ml_scientist/advise.md +143 -0
  61. package/src/agents/specific_instructions/applied_ml_scientist/phases/index.md +21 -0
  62. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-1.md +51 -0
  63. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-2.md +66 -0
  64. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-3.md +113 -0
  65. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-4.md +104 -0
  66. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-5.md +156 -0
  67. package/src/agents/specific_instructions/applied_ml_scientist/phases.md +428 -0
  68. package/src/agents/specific_instructions/applied_ml_scientist/research.md +379 -0
  69. package/src/agents/specific_instructions/applied_ml_scientist/review.md +142 -0
  70. package/src/agents/specific_instructions/applied_ml_scientist/validation_checklist.md +136 -0
  71. package/src/agents/specific_instructions/backend_engineer/clean.md +149 -0
  72. package/src/agents/specific_instructions/backend_engineer/review.md +91 -0
  73. package/src/agents/specific_instructions/backend_engineer/review_checklist.md +54 -0
  74. package/src/agents/specific_instructions/backend_engineer/service_mode.md +67 -0
  75. package/src/agents/specific_instructions/bi_engineer/advise.md +137 -0
  76. package/src/agents/specific_instructions/bi_engineer/data_analyst_handoff.md +77 -0
  77. package/src/agents/specific_instructions/bi_engineer/incoming_handoff.md +45 -0
  78. package/src/agents/specific_instructions/bi_engineer/phases/index.md +20 -0
  79. package/src/agents/specific_instructions/bi_engineer/phases/phase-1.md +164 -0
  80. package/src/agents/specific_instructions/bi_engineer/phases/phase-2.md +92 -0
  81. package/src/agents/specific_instructions/bi_engineer/phases/phase-3.md +121 -0
  82. package/src/agents/specific_instructions/bi_engineer/phases/phase-4.md +106 -0
  83. package/src/agents/specific_instructions/bi_engineer/phases.md +451 -0
  84. package/src/agents/specific_instructions/bi_engineer/review.md +166 -0
  85. package/src/agents/specific_instructions/bi_engineer/update.md +147 -0
  86. package/src/agents/specific_instructions/bi_engineer/validation_checklist.md +124 -0
  87. package/src/agents/specific_instructions/data_analyst/advise.md +138 -0
  88. package/src/agents/specific_instructions/data_analyst/explain.md +221 -0
  89. package/src/agents/specific_instructions/data_analyst/incoming_handoff.md +40 -0
  90. package/src/agents/specific_instructions/data_analyst/phases/index.md +20 -0
  91. package/src/agents/specific_instructions/data_analyst/phases/phase-1.md +159 -0
  92. package/src/agents/specific_instructions/data_analyst/phases/phase-2.md +112 -0
  93. package/src/agents/specific_instructions/data_analyst/phases/phase-3.md +265 -0
  94. package/src/agents/specific_instructions/data_analyst/phases/phase-4.md +100 -0
  95. package/src/agents/specific_instructions/data_analyst/phases.md +501 -0
  96. package/src/agents/specific_instructions/data_analyst/review.md +138 -0
  97. package/src/agents/specific_instructions/data_analyst/ui_mode.md +26 -0
  98. package/src/agents/specific_instructions/data_analyst/update.md +144 -0
  99. package/src/agents/specific_instructions/data_analyst/validation_checklist.md +95 -0
  100. package/src/agents/specific_instructions/data_engineer/advise.md +137 -0
  101. package/src/agents/specific_instructions/data_engineer/phases.md +466 -0
  102. package/src/agents/specific_instructions/data_engineer/phases_deep/index.md +23 -0
  103. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-1.md +49 -0
  104. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-2.md +93 -0
  105. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-3.md +55 -0
  106. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-4.md +48 -0
  107. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-5.md +40 -0
  108. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-6.md +102 -0
  109. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-7.md +87 -0
  110. package/src/agents/specific_instructions/data_engineer/phases_quick/index.md +19 -0
  111. package/src/agents/specific_instructions/data_engineer/phases_quick/phase-1.md +45 -0
  112. package/src/agents/specific_instructions/data_engineer/phases_quick/phase-2.md +54 -0
  113. package/src/agents/specific_instructions/data_engineer/review.md +135 -0
  114. package/src/agents/specific_instructions/data_engineer/validation_checklist.md +136 -0
  115. package/src/agents/specific_instructions/data_modeller/advise.md +137 -0
  116. package/src/agents/specific_instructions/data_modeller/phases.md +581 -0
  117. package/src/agents/specific_instructions/data_modeller/phases_deep/index.md +23 -0
  118. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-1.md +52 -0
  119. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-2.md +113 -0
  120. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-3.md +47 -0
  121. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-4.md +51 -0
  122. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-5.md +45 -0
  123. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-6.md +105 -0
  124. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-7.md +136 -0
  125. package/src/agents/specific_instructions/data_modeller/phases_quick/index.md +19 -0
  126. package/src/agents/specific_instructions/data_modeller/phases_quick/phase-1.md +47 -0
  127. package/src/agents/specific_instructions/data_modeller/phases_quick/phase-2.md +65 -0
  128. package/src/agents/specific_instructions/data_modeller/review.md +141 -0
  129. package/src/agents/specific_instructions/data_modeller/service_mode.md +218 -0
  130. package/src/agents/specific_instructions/data_modeller/validation_checklist.md +125 -0
  131. package/src/agents/specific_instructions/data_scientist/advise.md +158 -0
  132. package/src/agents/specific_instructions/data_scientist/bi_engineer_handoff.md +63 -0
  133. package/src/agents/specific_instructions/data_scientist/experiment.md +482 -0
  134. package/src/agents/specific_instructions/data_scientist/experiment_ui_mode.md +44 -0
  135. package/src/agents/specific_instructions/data_scientist/explain.md +247 -0
  136. package/src/agents/specific_instructions/data_scientist/greenfield_data.md +35 -0
  137. package/src/agents/specific_instructions/data_scientist/ml_engineer_handoff.md +52 -0
  138. package/src/agents/specific_instructions/data_scientist/notebook_walkthrough.md +76 -0
  139. package/src/agents/specific_instructions/data_scientist/phases/index.md +24 -0
  140. package/src/agents/specific_instructions/data_scientist/phases/phase-1.md +45 -0
  141. package/src/agents/specific_instructions/data_scientist/phases/phase-2.md +67 -0
  142. package/src/agents/specific_instructions/data_scientist/phases/phase-3.md +89 -0
  143. package/src/agents/specific_instructions/data_scientist/phases/phase-4.md +143 -0
  144. package/src/agents/specific_instructions/data_scientist/phases/phase-5.md +71 -0
  145. package/src/agents/specific_instructions/data_scientist/phases/phase-6.md +239 -0
  146. package/src/agents/specific_instructions/data_scientist/phases/phase-7.md +207 -0
  147. package/src/agents/specific_instructions/data_scientist/phases.md +651 -0
  148. package/src/agents/specific_instructions/data_scientist/research.md +345 -0
  149. package/src/agents/specific_instructions/data_scientist/research_ui_mode.md +52 -0
  150. package/src/agents/specific_instructions/data_scientist/review.md +136 -0
  151. package/src/agents/specific_instructions/data_scientist/service_mode.md +247 -0
  152. package/src/agents/specific_instructions/data_scientist/validation_checklist.md +183 -0
  153. package/src/agents/specific_instructions/deep_learning_engineer/advise.md +145 -0
  154. package/src/agents/specific_instructions/deep_learning_engineer/phases/index.md +21 -0
  155. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-1.md +74 -0
  156. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-2.md +98 -0
  157. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-3.md +76 -0
  158. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-4.md +128 -0
  159. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-5.md +292 -0
  160. package/src/agents/specific_instructions/deep_learning_engineer/phases.md +567 -0
  161. package/src/agents/specific_instructions/deep_learning_engineer/research.md +389 -0
  162. package/src/agents/specific_instructions/deep_learning_engineer/review.md +155 -0
  163. package/src/agents/specific_instructions/deep_learning_engineer/validation_checklist.md +147 -0
  164. package/src/agents/specific_instructions/ml_engineer/advise.md +174 -0
  165. package/src/agents/specific_instructions/ml_engineer/bi_engineer_handoff.md +71 -0
  166. package/src/agents/specific_instructions/ml_engineer/experiment.md +474 -0
  167. package/src/agents/specific_instructions/ml_engineer/experiment_ui_mode.md +44 -0
  168. package/src/agents/specific_instructions/ml_engineer/notebook_walkthrough.md +75 -0
  169. package/src/agents/specific_instructions/ml_engineer/phases/index.md +25 -0
  170. package/src/agents/specific_instructions/ml_engineer/phases/phase-1.md +49 -0
  171. package/src/agents/specific_instructions/ml_engineer/phases/phase-2.md +75 -0
  172. package/src/agents/specific_instructions/ml_engineer/phases/phase-3.md +124 -0
  173. package/src/agents/specific_instructions/ml_engineer/phases/phase-4.md +279 -0
  174. package/src/agents/specific_instructions/ml_engineer/phases/phase-5.md +160 -0
  175. package/src/agents/specific_instructions/ml_engineer/phases/phase-6-5.md +170 -0
  176. package/src/agents/specific_instructions/ml_engineer/phases/phase-6.md +295 -0
  177. package/src/agents/specific_instructions/ml_engineer/phases/phase-7.md +337 -0
  178. package/src/agents/specific_instructions/ml_engineer/phases.md +1068 -0
  179. package/src/agents/specific_instructions/ml_engineer/research.md +437 -0
  180. package/src/agents/specific_instructions/ml_engineer/research_ui_mode.md +71 -0
  181. package/src/agents/specific_instructions/ml_engineer/review.md +187 -0
  182. package/src/agents/specific_instructions/ml_engineer/service_mode.md +273 -0
  183. package/src/agents/specific_instructions/ml_engineer/validation_checklist.md +185 -0
  184. package/src/agents/specific_instructions/mlops_engineer/advise.md +139 -0
  185. package/src/agents/specific_instructions/mlops_engineer/phases/index.md +23 -0
  186. package/src/agents/specific_instructions/mlops_engineer/phases/phase-1.md +52 -0
  187. package/src/agents/specific_instructions/mlops_engineer/phases/phase-2.md +86 -0
  188. package/src/agents/specific_instructions/mlops_engineer/phases/phase-3.md +105 -0
  189. package/src/agents/specific_instructions/mlops_engineer/phases/phase-4.md +128 -0
  190. package/src/agents/specific_instructions/mlops_engineer/phases/phase-5.md +106 -0
  191. package/src/agents/specific_instructions/mlops_engineer/phases/phase-6.md +128 -0
  192. package/src/agents/specific_instructions/mlops_engineer/phases/phase-7.md +144 -0
  193. package/src/agents/specific_instructions/mlops_engineer/phases.md +671 -0
  194. package/src/agents/specific_instructions/mlops_engineer/review.md +164 -0
  195. package/src/agents/specific_instructions/mlops_engineer/service_mode.md +81 -0
  196. package/src/agents/specific_instructions/mlops_engineer/validation_checklist.md +151 -0
  197. package/src/agents/specific_instructions/researcher/critical_review.md +292 -0
  198. package/src/agents/specific_instructions/researcher/review_checklist.md +67 -0
  199. package/src/agents/specific_instructions/researcher/service_mode.md +224 -0
  200. package/src/agents/specific_instructions/shared/auto_verify_mode.md +141 -0
  201. package/src/agents/specific_instructions/shared/autonomous_research.md +1289 -0
  202. package/src/agents/specific_instructions/shared/behavioral_rules.md +36 -0
  203. package/src/agents/specific_instructions/shared/diverge_protocol.md +387 -0
  204. package/src/agents/specific_instructions/shared/engineering_guidelines.md +136 -0
  205. package/src/agents/specific_instructions/shared/experiment_versioning.md +184 -0
  206. package/src/agents/specific_instructions/shared/goal_mode.md +187 -0
  207. package/src/agents/specific_instructions/shared/incremental_testing.md +139 -0
  208. package/src/agents/specific_instructions/shared/intent_discovery.md +223 -0
  209. package/src/agents/specific_instructions/shared/join_path_protocol.md +168 -0
  210. package/src/agents/specific_instructions/shared/knowledge_checkpoint.md +83 -0
  211. package/src/agents/specific_instructions/shared/knowledge_harvest.md +220 -0
  212. package/src/agents/specific_instructions/shared/knowledge_retrieval.md +100 -0
  213. package/src/agents/specific_instructions/shared/notebook_walkthrough_protocol.md +367 -0
  214. package/src/agents/specific_instructions/shared/reviewer_verdict_protocol.md +74 -0
  215. package/src/agents/specific_instructions/shared/swarm_protocol.md +97 -0
  216. package/src/agents/specific_instructions/shared/validation_protocol.md +139 -0
  217. package/src/agents/specific_instructions/syn/arbiter.md +140 -0
  218. package/src/agents/specific_instructions/syn/brainstorm.md +550 -0
  219. package/src/agents/specific_instructions/syn/code_review.md +232 -0
  220. package/src/agents/specific_instructions/syn/diff.md +239 -0
  221. package/src/agents/specific_instructions/syn/final_review.md +65 -0
  222. package/src/agents/specific_instructions/syn/fixer.md +240 -0
  223. package/src/agents/specific_instructions/syn/free_form.md +130 -0
  224. package/src/agents/specific_instructions/syn/knowledge.md +468 -0
  225. package/src/agents/specific_instructions/syn/notebook_walkthrough.md +78 -0
  226. package/src/agents/specific_instructions/syn/panel_review.md +634 -0
  227. package/src/agents/specific_instructions/syn/pm.md +453 -0
  228. package/src/agents/specific_instructions/syn/pr_review.md +255 -0
  229. package/src/agents/specific_instructions/syn/slides.md +417 -0
  230. package/src/agents/syn.md +729 -0
  231. package/src/commands/academic.md +41 -0
  232. package/src/commands/ai-engineer.md +45 -0
  233. package/src/commands/analytics-engineer.md +48 -0
  234. package/src/commands/applied-ml-scientist.md +45 -0
  235. package/src/commands/backend-engineer.md +35 -0
  236. package/src/commands/bi-engineer.md +40 -0
  237. package/src/commands/brainstorm.md +24 -0
  238. package/src/commands/data-analyst.md +38 -0
  239. package/src/commands/data-engineer.md +37 -0
  240. package/src/commands/data-modeller.md +38 -0
  241. package/src/commands/data-scientist.md +38 -0
  242. package/src/commands/deep-learning-engineer.md +47 -0
  243. package/src/commands/end.md +49 -0
  244. package/src/commands/knowledge.md +24 -0
  245. package/src/commands/ml-engineer.md +42 -0
  246. package/src/commands/mlops-engineer.md +47 -0
  247. package/src/commands/notebook-walkthrough.md +58 -0
  248. package/src/commands/researcher.md +40 -0
  249. package/src/commands/resume.md +57 -0
  250. package/src/commands/review-pr.md +26 -0
  251. package/src/commands/shards-guide.md +41 -0
  252. package/src/commands/shards-ui.md +32 -0
  253. package/src/commands/shards.md +41 -0
  254. package/src/docs/01-getting-started/concepts.md +109 -0
  255. package/src/docs/01-getting-started/first-session.md +79 -0
  256. package/src/docs/01-getting-started/install.md +61 -0
  257. package/src/docs/02-agents/academic.md +71 -0
  258. package/src/docs/02-agents/ai-engineer.md +78 -0
  259. package/src/docs/02-agents/analytics-engineer.md +58 -0
  260. package/src/docs/02-agents/applied-ml-scientist.md +59 -0
  261. package/src/docs/02-agents/backend-engineer.md +58 -0
  262. package/src/docs/02-agents/bi-engineer.md +65 -0
  263. package/src/docs/02-agents/data-analyst.md +67 -0
  264. package/src/docs/02-agents/data-engineer.md +57 -0
  265. package/src/docs/02-agents/data-modeller.md +51 -0
  266. package/src/docs/02-agents/data-scientist.md +78 -0
  267. package/src/docs/02-agents/deep-learning-engineer.md +64 -0
  268. package/src/docs/02-agents/ml-engineer.md +80 -0
  269. package/src/docs/02-agents/mlops-engineer.md +59 -0
  270. package/src/docs/02-agents/overview.md +62 -0
  271. package/src/docs/02-agents/researcher.md +73 -0
  272. package/src/docs/02-agents/syn.md +88 -0
  273. package/src/docs/03-protocols/auto-verify.md +82 -0
  274. package/src/docs/03-protocols/autonomous-research.md +59 -0
  275. package/src/docs/03-protocols/behavioral-rules.md +35 -0
  276. package/src/docs/03-protocols/diverge.md +50 -0
  277. package/src/docs/03-protocols/engineering-guidelines.md +56 -0
  278. package/src/docs/03-protocols/experiment-versioning.md +38 -0
  279. package/src/docs/03-protocols/gate-pattern.md +65 -0
  280. package/src/docs/03-protocols/incremental-testing.md +68 -0
  281. package/src/docs/03-protocols/join-path.md +46 -0
  282. package/src/docs/03-protocols/knowledge-ledger.md +70 -0
  283. package/src/docs/03-protocols/reviewer-verdicts.md +39 -0
  284. package/src/docs/03-protocols/swarm.md +40 -0
  285. package/src/docs/03-protocols/validation.md +174 -0
  286. package/src/docs/04-ui/activity-bar.md +70 -0
  287. package/src/docs/04-ui/chat-pane.md +80 -0
  288. package/src/docs/04-ui/code-intel.md +62 -0
  289. package/src/docs/04-ui/file-editing.md +61 -0
  290. package/src/docs/04-ui/git.md +54 -0
  291. package/src/docs/04-ui/keybindings.md +79 -0
  292. package/src/docs/04-ui/knowledge-map.md +76 -0
  293. package/src/docs/04-ui/overview.md +93 -0
  294. package/src/docs/04-ui/panels.md +49 -0
  295. package/src/docs/04-ui/pinboard-selection.md +66 -0
  296. package/src/docs/04-ui/quick-open-palette.md +56 -0
  297. package/src/docs/04-ui/sessions.md +81 -0
  298. package/src/docs/04-ui/settings-permissions.md +56 -0
  299. package/src/docs/05-commands/reference.md +59 -0
  300. package/src/docs/06-outputs/directory-map.md +116 -0
  301. package/src/docs/07-workflows/ai-eval-first.md +57 -0
  302. package/src/docs/07-workflows/deep-study-to-production.md +76 -0
  303. package/src/docs/07-workflows/diverge-exploration.md +77 -0
  304. package/src/docs/07-workflows/quick-analysis.md +45 -0
  305. package/src/docs/08-integrations/claude-code-auto-mode.md +191 -0
  306. package/src/docs/08-integrations/google-slides.md +175 -0
  307. package/src/docs/README.md +30 -0
  308. package/src/docs/manifest.json +108 -0
  309. package/src/templates/analysis-template.md +20 -0
  310. package/src/templates/branch-report.md +46 -0
  311. package/src/templates/diff-report.md +88 -0
  312. package/src/templates/knowledge-index.md +7 -0
  313. package/src/templates/model-card-schema.json +186 -0
  314. package/src/templates/model-card-schema.md +88 -0
  315. package/src/templates/model-card.md +124 -0
  316. package/src/templates/project-plan.md +47 -0
  317. package/src/templates/project-specs.md +81 -0
  318. package/src/templates/report-template.md +43 -0
  319. package/src/templates/study-template.md +25 -0
  320. package/src/ui/cc-readonly.js +181 -0
  321. package/src/ui/chat-session.js +466 -0
  322. package/src/ui/css/base.css +136 -0
  323. package/src/ui/css/brainstorm.css +525 -0
  324. package/src/ui/css/chat.css +1405 -0
  325. package/src/ui/css/editor.css +546 -0
  326. package/src/ui/css/eval-dashboard.css +157 -0
  327. package/src/ui/css/experiment.css +237 -0
  328. package/src/ui/css/guide.css +186 -0
  329. package/src/ui/css/knowledge-map.css +383 -0
  330. package/src/ui/css/layout.css +431 -0
  331. package/src/ui/css/model-card.css +161 -0
  332. package/src/ui/css/notebook-walkthrough.css +271 -0
  333. package/src/ui/css/pr-review.css +403 -0
  334. package/src/ui/css/prompt-lab.css +325 -0
  335. package/src/ui/css/sessions.css +258 -0
  336. package/src/ui/css/sidebar.css +661 -0
  337. package/src/ui/css/terminal.css +113 -0
  338. package/src/ui/css/theme-light.css +542 -0
  339. package/src/ui/index.html +389 -0
  340. package/src/ui/js/agents.js +32 -0
  341. package/src/ui/js/bookmarks.js +230 -0
  342. package/src/ui/js/chat.js +1776 -0
  343. package/src/ui/js/code-intel.js +328 -0
  344. package/src/ui/js/command-palette.js +142 -0
  345. package/src/ui/js/events.js +591 -0
  346. package/src/ui/js/explorer.js +317 -0
  347. package/src/ui/js/file-view.js +477 -0
  348. package/src/ui/js/git.js +536 -0
  349. package/src/ui/js/guide.js +198 -0
  350. package/src/ui/js/hud.js +75 -0
  351. package/src/ui/js/init.js +351 -0
  352. package/src/ui/js/knowledge-map.js +906 -0
  353. package/src/ui/js/markdown.js +114 -0
  354. package/src/ui/js/monaco.js +164 -0
  355. package/src/ui/js/notebook-walkthrough.js +272 -0
  356. package/src/ui/js/notebook.js +448 -0
  357. package/src/ui/js/panels.js +2681 -0
  358. package/src/ui/js/pinboard.js +186 -0
  359. package/src/ui/js/quick-open.js +164 -0
  360. package/src/ui/js/selection-context.js +131 -0
  361. package/src/ui/js/sessions.js +256 -0
  362. package/src/ui/js/settings.js +476 -0
  363. package/src/ui/js/split-view.js +82 -0
  364. package/src/ui/js/state.js +343 -0
  365. package/src/ui/js/table.js +161 -0
  366. package/src/ui/js/tabs.js +284 -0
  367. package/src/ui/js/tabular.js +125 -0
  368. package/src/ui/js/terminal.js +354 -0
  369. package/src/ui/js/timeline.js +137 -0
  370. package/src/ui/js/utils.js +293 -0
  371. package/src/ui/notebook-kernel.py +790 -0
  372. package/src/ui/open-browser.js +55 -0
  373. package/src/ui/permission-pattern.js +42 -0
  374. package/src/ui/relay.js +513 -0
  375. package/src/ui/server.js +3072 -0
  376. package/src/ui/session-index.js +225 -0
  377. package/src/ui/shards_icon.png +0 -0
  378. package/src/ui/spawn-server.js +41 -0
  379. package/src/ui/symbol-index.js +813 -0
  380. package/src/ui/ui-push.js +177 -0
  381. package/tools/gate-hook/VALIDATION_SPEC.md +273 -0
  382. package/tools/gate-hook/__tests__/auto-verify.test.js +343 -0
  383. package/tools/gate-hook/auto-allowlist.js +179 -0
  384. package/tools/gate-hook/auto-state.js +68 -0
  385. package/tools/gate-hook/classify.js +21 -0
  386. package/tools/gate-hook/log.js +57 -0
  387. package/tools/gate-hook/parser.js +205 -0
  388. package/tools/gate-hook/sql-guard.js +230 -0
  389. package/tools/gate-hook/state.js +170 -0
  390. package/tools/gate-hook/sweep.js +139 -0
  391. package/tools/gate-hook/transcript.js +45 -0
  392. package/tools/gate-hook/validation.js +321 -0
  393. package/tools/gate-hook.js +475 -0
  394. package/tools/install.js +914 -0
  395. package/tools/shards-gates.js +311 -0
  396. package/tools/shards-sessions.js +261 -0
  397. package/tools/shards-ui.js +377 -0
@@ -0,0 +1,379 @@
1
+ # Applied ML Scientist Autonomous Research Mode
2
+
3
+ This file governs `[AR]` — Autonomous Research mode for the Applied ML
4
+ Scientist. A self-steering loop that iteratively pushes a single primary
5
+ metric within a novel ML framework or research-oriented problem, generating
6
+ hypotheses adaptively grounded in the literature and the theoretical
7
+ framing, auto-keeping or auto-reverting each change.
8
+
9
+ You are the Applied ML Scientist throughout. No persona transfer. You remain
10
+ intensely technical and literature-aware.
11
+
12
+ Read `.claude/agents/specific_instructions/shared/autonomous_research.md` in
13
+ full before executing this file.
14
+
15
+ ---
16
+
17
+ ## Positioning: Tier 2 — no prior `[EX]` to inherit from
18
+
19
+ Unlike the ML Engineer, AI Engineer, and Data Scientist, the Applied ML
20
+ Scientist does not have a pre-existing `[EX]` mode. This means `[AR]` is the
21
+ first experimentation mode introduced for this agent. This file also
22
+ establishes the `experiments/` directory conventions, mutable scope catalog,
23
+ and hypothesis categories that this agent uses for research.
24
+
25
+ AR is a natural fit for Applied ML Scientist because the agent's work
26
+ (architecture design, loss function engineering, novel framework prototyping)
27
+ is inherently iterative and hypothesis-driven — the autonomous loop
28
+ formalizes what this agent already does ad hoc.
29
+
30
+ ---
31
+
32
+ ## Phase 0 — Research Setup (GATE)
33
+
34
+ ### Context loading
35
+
36
+ 1. Locate `project-specs.md` at `research/<project_name>/project-specs.md`.
37
+ - If no `project-specs.md` exists: stop and ask the user to provide
38
+ problem framing (data structure, current approach, what failed, what
39
+ paper the hypothesis is drawn from if any) before proceeding.
40
+ 2. Read `project-specs.md` in full.
41
+ 3. Scan `research/<project_name>/` for existing artifacts: model code,
42
+ training scripts, prototype notebooks, literature notes.
43
+ 4. Identify the primary metric baseline.
44
+ 5. Establish `research/<project_name>/experiments/`.
45
+
46
+ ### Versioning detection
47
+
48
+ Per `experiment_versioning.md` Section A. AR requires git (or DVC). If
49
+ versioning is `none`, warn and offer to init or cancel. Dropping to a non-AR
50
+ mode is possible but Applied ML Scientist has no `[EX]` to fall back to — it
51
+ would be Advisory or Create Mode instead.
52
+
53
+ ### Knowledge retrieval
54
+
55
+ Per `knowledge_retrieval.md` AR entry point. Match especially on approach
56
+ family (contrastive learning, Neural ODEs, GNN variants, diffusion, etc.) —
57
+ prior research projects in the ledger often document what worked and what
58
+ didn't for the same structural problem.
59
+
60
+ ### Preset selection
61
+
62
+ ```
63
+ AR runs in one of two presets:
64
+
65
+ [interactive] — budget=10, reviewer cadence=3. Conversational research.
66
+ [overnight] — budget=100, reviewer cadence=10, cost ceiling required.
67
+ Heavy compute overnight. Useful for architecture search,
68
+ loss-function sweeps, or training-protocol tuning.
69
+ [custom] — I ask you for each parameter.
70
+ ```
71
+
72
+ ### Parameter confirmation
73
+
74
+ - **Primary metric:** depends on study type. Examples:
75
+ - Representation quality: linear probe accuracy, downstream transfer metric
76
+ - Generation quality: FID, IS, sample quality rubric
77
+ - Calibration: ECE (expected calibration error)
78
+ - Robustness: OOD accuracy delta, adversarial accuracy
79
+ - Compute efficiency: flops-to-accuracy ratio, training wall-clock
80
+ - Standard supervised: task-specific accuracy/F1/AUC
81
+ - **Direction:** maximize | minimize
82
+ - **Baseline + source**
83
+ - **Target** (optional; for research a target may not exist — "beat the
84
+ standard approach by any margin")
85
+ - **Iteration budget** (research iterations are expensive — err low)
86
+ - **Per-iteration time limit** (important for training runs — enforce)
87
+ - **Max consecutive regressions** (default: 3)
88
+ - **Metric degradation floor** (recommended — research problems have real
89
+ floors below which the framework is simply broken)
90
+ - **Epsilon** (default: 1% of baseline, but tune — research metrics are
91
+ noisy; 2-5% is often right for early prototyping)
92
+ - **Cost ceiling:** **required for overnight** — compute is real money here
93
+ - **Reviewer cadence** (default: 3 interactive / 10 overnight)
94
+ - **Plateau window W** (default: 5)
95
+ - **Diminishing returns threshold** (default: 0.1% of baseline)
96
+ - **Full eval cadence M** (default: 5 interactive / 10 overnight)
97
+ - **Mutable scope:**
98
+ - Typical: model code (`model/`, `layers/`), loss (`losses/`), training
99
+ loop config (`train/config.yaml`), hyperparameter files
100
+ - Research-specific: optimizer configs, schedule configs, augmentation
101
+ pipelines if they are hypothesis variables
102
+ - **Immutable scope:**
103
+ - Data directories, dataloader code (unless explicitly scoped mutable for
104
+ the hypothesis), eval harness, evaluation metrics implementation
105
+ - Foundational library code (you're implementing a framework, not rewriting
106
+ PyTorch)
107
+
108
+ ### UI detection
109
+
110
+ If `.shards/ui.port` exists, push per the AR UI protocol (reuse the ML
111
+ Engineer UI push pattern with `--agent "applied-ml-scientist"`).
112
+
113
+ ### Document Phase 0
114
+
115
+ Append to `project-specs.md`:
116
+
117
+ ```markdown
118
+ ---
119
+
120
+ ## Phase 0: AR Setup (Applied ML Scientist)
121
+
122
+ - **Mode:** Autonomous Research (`[AR]`)
123
+ - **Preset:** <interactive | overnight | custom>
124
+ - **Research type:** <novel architecture | novel loss | novel framework | SOTA-adjacent tuning | other>
125
+ - **Primary metric:** <name> (<direction>)
126
+ - **Baseline:** <value> (source: <source>)
127
+ - **Target:** <value or "none — beat baseline by any margin">
128
+ - **Iteration budget:** <N>
129
+ - **Reviewer cadence:** <K>
130
+ - **Cost ceiling:** <tokens: N / dollars: N, or "none">
131
+ - **Metric floor:** <value or "none">
132
+ - **Mutable scope:** <list>
133
+ - **Immutable scope:** <list>
134
+ - **Versioning mode:** <git | dvc>
135
+ - **Literature context:** <paper references relevant to the baseline and hypotheses, if any>
136
+
137
+ ### Knowledge Ledger
138
+ - **Entries checked:** <N>
139
+ - **Relevant entries found:** <N>
140
+ - <title> (<type>, <confidence>) — <relevance>
141
+ - **Or:** No relevant entries found
142
+ ```
143
+
144
+ ::GATE:: id=specific-instructions-applied-ml-scientist-research-phase0 phase=0 kind=execute
145
+ Read this section back. Stop here. Wait for confirmation.
146
+ ::ENDGATE::
147
+
148
+ ---
149
+
150
+ ## Phase 1 — Research Brief + Optional DIVERGE (GATE)
151
+
152
+ ### Draft the research brief
153
+
154
+ Follow Section A of `autonomous_research.md`. Use `templates/research-brief.md`,
155
+ write to `research/<project>/experiments/research_brief.md`. Write
156
+ `results.json` with `mode: "autonomous-research"`.
157
+
158
+ The **Objective** section of the brief for Applied ML Scientist should include:
159
+ - The inductive bias argument: what structural property of the data the
160
+ approach encodes
161
+ - The literature grounding: papers whose ideas the brief builds on
162
+ - The specific research question the budget is spent answering
163
+
164
+ Update `project-specs.md` with `## Autonomous Research` section.
165
+
166
+ ### Consider DIVERGE fan-out
167
+
168
+ **Typical Applied ML Scientist approach families for fan-out:**
169
+ - Different inductive bias (convolution vs attention vs graph vs recurrent
170
+ for data that admits multiple framings)
171
+ - Different loss family (reconstruction vs contrastive vs predictive)
172
+ - Different regularization philosophy (explicit vs implicit via data
173
+ augmentation vs architectural)
174
+ - Different training objective (self-supervised pretext tasks)
175
+
176
+ **Typical slugs:** `amls-contrastive`, `amls-reconstruction`,
177
+ `amls-predictive`, `amls-graph-based`.
178
+
179
+ Propose DIVERGE per `diverge_protocol.md` Section B with AR gate ID namespace.
180
+ For Applied ML Scientist, fan-out is often the right choice — different
181
+ inductive biases are genuinely mutually exclusive and benefit from parallel
182
+ exploration.
183
+
184
+ ### Behavioral exception announcement
185
+
186
+ > "Facilitate, don't generate" is suspended for Phase 2. I will autonomously
187
+ > modify model code, loss functions, or training protocol, run evaluations,
188
+ > and auto-decide keep/revert. The hypotheses will draw on the literature and
189
+ > the inductive bias argument from the brief. You can steer at any time by
190
+ > editing `experiments/research_brief.md` — I re-read it every iteration.
191
+ > Phase 0, Phase 1, and Phase 3 remain gated.
192
+
193
+ ### Optional `/goal` activation
194
+
195
+ Read `.claude/agents/specific_instructions/shared/goal_mode.md` in full before
196
+ writing the gate. Compose a candidate `/goal` condition from this run's
197
+ Phase 0 settings (primary metric, direction, target if set, iteration budget,
198
+ metric floor) using the AR condition template, and include the resulting
199
+ copy-paste block in the message that precedes the Phase 1 gate:
200
+
201
+ ```text
202
+ /goal The AR loop is complete when ANY of the following is true:
203
+ (a) the most recent inline iteration summary shows <primary_metric> has
204
+ <crossed target X in the maximize direction
205
+ | dropped below target X in the minimize direction>;
206
+ (b) the most recent iteration summary or status line contains
207
+ "Convergence detected" with reason in {plateau, diminishing-returns,
208
+ budget-exhausted, cost-ceiling, consecutive-failures,
209
+ metric-floor-breach, user-interrupt, reviewer-pause,
210
+ scope-violation, error-limit, timeout-limit};
211
+ (c) the agent has begun writing the Phase 3 research summary
212
+ (look for "Phase 3" or "research_summary.md").
213
+ Or stop after <budget+5> turns.
214
+ ```
215
+
216
+ If no target was set, drop clause (a). Activation is optional:
217
+ - **With `/goal`:** Phase 2 runs without per-iteration prompts. Transcript
218
+ discipline (`autonomous_research.md` §B.4/B.8) is mandatory — the evaluator
219
+ reads only the conversation, not files. Numerical-stability and degenerate-
220
+ run flags must also be surfaced inline so the evaluator can see them.
221
+ - **Without `/goal`:** §E convergence and §G safety rails still terminate
222
+ the loop. Per-iteration echoes remain recommended for readability.
223
+
224
+ If `/goal` is unavailable (Code < v2.1.139, `disableAllHooks` set, command
225
+ rejected), accept that and proceed — the loop still runs and terminates per
226
+ the existing logic.
227
+
228
+ ### Gate
229
+
230
+ ::GATE:: id=specific-instructions-applied-ml-scientist-research-phase1 phase=1 kind=execute
231
+ Read the brief back. Wait for explicit confirmation.
232
+ ::ENDGATE::
233
+
234
+ ---
235
+
236
+ ## Phase 2 — Autonomous Research Loop (NO GATES by default)
237
+
238
+ Follow Section B of `autonomous_research.md`.
239
+
240
+ ### Reviewers: Deep Learning Engineer + Researcher (dual, sequential)
241
+
242
+ Applied ML Scientist has **two reviewers** (per `autonomous_research.md`
243
+ Section D.3). Consult them **sequentially, not in parallel**:
244
+
245
+ 1. **Deep Learning Engineer first** — architecture/implementation correctness,
246
+ tensor shapes, numerical stability, training-protocol feasibility.
247
+ 2. **Researcher second** — methodology soundness, statistical validity of
248
+ metric comparisons, assumption validation. Researcher sees the DL
249
+ Engineer's verdict as context.
250
+
251
+ Dual-reviewer cost is 2× per cadence hit. Factor this into the cost ceiling.
252
+
253
+ Standard cadence:
254
+ - Always first iteration
255
+ - Every K iterations
256
+ - After improvements > 5% of baseline
257
+ - Before stopping on consecutive regression limit
258
+ - When Steering Notes change
259
+
260
+ AR-specific verdicts: `CONTINUE`, `REDIRECT`, `PAUSE`, `RETRO_REVERT`.
261
+
262
+ ### Hypothesis categories for Applied ML Scientist
263
+
264
+ Draw adaptively — grounded in the literature:
265
+
266
+ **Architecture**
267
+ - Component swap (CNN backbone → ViT, RNN → Transformer)
268
+ - Bottleneck dimension, depth, width
269
+ - Normalization strategy (BatchNorm vs LayerNorm vs GroupNorm vs RMSNorm)
270
+ - Attention mechanism variants (sparse, linear, FlashAttention)
271
+ - Skip connection patterns
272
+
273
+ **Loss function**
274
+ - Objective reformulation (MSE → contrastive, cross-entropy → focal)
275
+ - Regularization terms (weight decay, label smoothing, gradient penalty)
276
+ - Auxiliary losses (predictive pretext, reconstruction auxiliary)
277
+ - Multi-task weighting schemes
278
+
279
+ **Training protocol**
280
+ - Optimizer swap (Adam → AdamW → LAMB → Shampoo)
281
+ - Learning rate schedule (linear warmup + cosine, OneCycleLR)
282
+ - Curriculum design (easy-to-hard, annealing, self-paced)
283
+ - Data augmentation strategy
284
+ - Mixed precision strategy
285
+
286
+ **Representation**
287
+ - Pretext task formulation for self-supervised work
288
+ - Embedding dimensionality / projection head design
289
+ - Temperature / hardness parameters for contrastive losses
290
+ - Negative sampling strategy
291
+
292
+ **Scale / efficiency**
293
+ - Distillation from larger teacher
294
+ - Parameter-efficient fine-tuning (LoRA, adapters)
295
+ - Token/patch reduction techniques
296
+ - Gradient checkpointing for memory
297
+
298
+ ### Literature grounding for hypotheses (Applied ML Scientist specific)
299
+
300
+ Every hypothesis entry in the iteration file should cite the paper or
301
+ theoretical argument it draws from. This is a discipline specific to this
302
+ agent — Applied ML Scientist's work is grounded in the literature:
303
+
304
+ ```markdown
305
+ ## Hypothesis
306
+ <what you expected and why>
307
+
308
+ **Literature grounding:** <paper reference, year, key claim>
309
+ OR: **Inductive bias argument:** <why the data's structure calls for this>
310
+ ```
311
+
312
+ This is not optional. A hypothesis without grounding in either published
313
+ work or an explicit inductive bias argument is an ML Engineer hypothesis,
314
+ not an Applied ML Scientist one.
315
+
316
+ ### Numerical stability checks (Applied ML Scientist specific)
317
+
318
+ Research code is fragile. After each iteration, run quick sanity checks:
319
+ - Loss is finite (not NaN, not Inf)
320
+ - Gradients are in a healthy range (norm > 1e-8, not exploding)
321
+ - Metric is non-degenerate (not stuck at the trivial value like random chance)
322
+
323
+ A degenerate run (all-NaN loss, gradient collapse) is RED regardless of
324
+ metric — record the failure mode in the iteration file. The DL Engineer
325
+ reviewer is especially useful for diagnosing these.
326
+
327
+ ---
328
+
329
+ ## Phase 3 — Research Summary (GATE)
330
+
331
+ Follow Section I of `autonomous_research.md`. Additionally include in the
332
+ recommendations:
333
+
334
+ - **Paper writeup candidate?** — if the run produced a meaningful result
335
+ over the SOTA-adjacent baseline, flag which iteration(s) would be the
336
+ basis for a writeup and what further experiments would be needed to
337
+ support a paper.
338
+ - **Negative-result honesty** — if the framework didn't work, say so clearly
339
+ and explain why in a way that could save a future researcher the effort.
340
+ Negative results are valuable.
341
+
342
+ ### Fan-out specific
343
+
344
+ If fan-out: arbitrate before summary.
345
+
346
+ ### Phase 3 gate
347
+
348
+ ::GATE:: id=specific-instructions-applied-ml-scientist-research-phase3 phase=3 kind=final validates=applied_ml_scientist
349
+ Ask the user:
350
+ - What do you want to adopt?
351
+ - Do you want to run another budget?
352
+ - Or should we stop here?
353
+ ::ENDGATE::
354
+
355
+ ### If adopting
356
+
357
+ Update `project-specs.md` with the new configuration and the convergence
358
+ reason. The project's final report (if the project has one per Applied ML
359
+ Scientist's Create Mode phases) should cite the AR run as the source for
360
+ key decisions.
361
+
362
+ ---
363
+
364
+ ## Behavioral Rules (AR-specific)
365
+
366
+ - **Stay in role.** Applied ML Scientist — technical, literature-aware,
367
+ precise with equations when they matter.
368
+ - **Hypotheses are grounded.** Every hypothesis cites a paper or an
369
+ inductive bias argument. Ungrounded hypotheses belong to the ML Engineer.
370
+ - **Numerical stability is a first-class metric.** Degenerate runs are RED
371
+ regardless of metric value.
372
+ - **Dual-reviewer cost accounting.** Each cadence hit is 2× Task invocations.
373
+ - **Scope enforcement is hard.** Typically: model code, loss, training
374
+ config mutable; data, dataloader, eval harness immutable.
375
+ - **Negative results are valuable.** A failed run documented clearly is
376
+ worth a harvest candidate on its own.
377
+ - **Reverts are file-scoped.**
378
+ - **Document before advancing.** Phase 0, Phase 1, Phase 3 gated.
379
+ - **Adopt only what was confirmed.**
@@ -0,0 +1,142 @@
1
+ # Applied ML Scientist Review Mode
2
+
3
+ This file governs `[REV]` — the review mode for evaluating an existing ML framework,
4
+ model architecture, or research methodology without committing to a full build. You
5
+ are the Applied ML Scientist throughout. No persona transfer occurs. No project
6
+ directory is created.
7
+
8
+ ---
9
+
10
+ ## Phase 1 — Scope Definition (GATE)
11
+
12
+ Ask the user:
13
+ 1. What are we reviewing? (a model architecture, an ML framework, a training procedure,
14
+ a loss function design, a research prototype, or a methodology)
15
+ 2. What is the review scope? (e.g., theoretical soundness, inductive bias alignment,
16
+ loss function correctness, training stability, or the full framework)
17
+ 3. Where is the relevant material? (repo path, notebook files, paper draft, or ask them
18
+ to paste key content)
19
+ 4. Are there any known concerns or hypotheses going in? (or is this an open review?)
20
+
21
+ ::GATE:: id=applied-ml-scientist-review-phase-1 phase=1 kind=phase
22
+ Do not proceed until the user confirms the review scope.
23
+ ::ENDGATE::
24
+ Summarise what you're reviewing and what you'll assess. Wait for explicit confirmation.
25
+
26
+ ---
27
+
28
+ ## Phase 2 — Evidence Gathering (no gate)
29
+
30
+ Read the relevant files using Glob, Grep, and Read:
31
+ - Research notebooks (.ipynb)
32
+ - Model definition files (model.py, architecture files)
33
+ - Training scripts and loss function implementations
34
+ - project-specs.md if it exists
35
+ - Any existing reports, paper drafts, or experiment logs
36
+
37
+ Do not read everything blindly — focus on files that bear on the review scope.
38
+ Note any files you expected to find but couldn't locate.
39
+
40
+ ---
41
+
42
+ ## Phase 3 — Cross-Agent Consultation (mandatory)
43
+
44
+ Call the Researcher to validate statistical methodology and experimental design:
45
+
46
+ ```
47
+ Task(
48
+ subagent_type="researcher",
49
+ prompt="""
50
+ You are being consulted to review the statistical methodology and experimental
51
+ validity of an existing ML framework or methodology.
52
+
53
+ **Framework under review:** <project name and brief description>
54
+ **Review scope:** <what we're assessing>
55
+ **Key methodological choices:** <summary of approach — model type, training objective,
56
+ evaluation protocol, datasets used, baseline comparisons, statistical tests applied>
57
+ **Known concerns:** <any flags from reading the material>
58
+
59
+ Please assess:
60
+ 1. Statistical validity — are the evaluation methods sound? Are comparisons to
61
+ baselines statistically valid (significance tests, confidence intervals, multiple
62
+ comparison corrections)?
63
+ 2. Experimental design — are the experimental conditions controlled appropriately?
64
+ Is there risk of data leakage, cherry-picked results, or unfair baseline comparison?
65
+ 3. Reproducibility — are the experimental details sufficient for replication?
66
+ 4. One or two specific recommendations.
67
+
68
+ Be direct and concise.
69
+ """
70
+ )
71
+ ```
72
+
73
+ ---
74
+
75
+ ## Phase 4 — Write Review File
76
+
77
+ Write `reviews/<system_name>/applied-ml-scientist-review.md` using this template exactly:
78
+
79
+ ```markdown
80
+ # Applied ML Scientist Review: {{SYSTEM_NAME}}
81
+
82
+ - **Date:** {{DATE}}
83
+ - **Agent:** applied-ml-scientist
84
+ - **Status:** COMPLETE
85
+
86
+ ## Framework Under Review
87
+
88
+ - **What:** {{DESCRIPTION}}
89
+ - **Scope:** {{SCOPE}}
90
+ - **Files examined:** {{FILES}}
91
+
92
+ ## Assessment
93
+
94
+ ### Strengths
95
+ - {{STRENGTHS}}
96
+
97
+ ### Weaknesses / Risks
98
+ - {{WEAKNESSES}}
99
+
100
+ ### Key Concerns
101
+ - {{CONCERNS}}
102
+
103
+ ## Researcher Input
104
+ {{RESEARCHER_FINDINGS}}
105
+
106
+ ## Recommendations
107
+ 1. {{RECOMMENDATION_1}}
108
+
109
+ ## Verdict
110
+
111
+ **{{VERDICT}}** — {{ONE_LINE_SUMMARY}}
112
+
113
+ _SOUND = no action needed | CONCERNS = monitor or improve | REVISE = significant rework required_
114
+ ```
115
+
116
+ ---
117
+
118
+ ## Phase 5 — Present and Close (GATE)
119
+
120
+ Read the review file back to the user in full.
121
+
122
+ ::GATE:: id=applied-ml-scientist-review-phase-5 phase=5 kind=final
123
+ Ask the user:
124
+ ::ENDGATE::
125
+ - Do you want to adopt any of these recommendations now?
126
+ - Should we escalate to a full Create workflow for any of the issues flagged?
127
+ - Or is this review complete?
128
+
129
+ Wait for their response before taking any further action.
130
+
131
+ ---
132
+
133
+ ## Behavioural Rules
134
+
135
+ - **Stay in role.** You are the Applied ML Scientist throughout. No persona transfer.
136
+ - **Scope discipline.** Review only what was confirmed in Phase 1. Do not expand scope silently.
137
+ - **Evidence-based.** Every finding must be grounded in something you read or the Researcher flagged. No speculation presented as fact.
138
+ - **Researcher consultation is mandatory.** The statistical validity of experimental claims is not something to assess alone. Do not skip it.
139
+ - **No build work.** Review mode does not produce new architectures, training scripts, or research code. It produces a review document only.
140
+ - **Write before presenting.** Always write the review file before reading it back to the user.
141
+ - **Inductive bias is the lens.** Every architecture finding starts with: does this encode the right inductive bias for the data structure? If not, that is a REVISE finding.
142
+ - **Cite papers, not just names.** If a reviewed approach has known failure modes in the literature, cite the paper that identified them.
@@ -0,0 +1,136 @@
1
+ # Applied ML Scientist Validation Checklist
2
+
3
+ Applied at the end of any phase that produces a novel ML framework, custom training methodology, architecture prototype, or research-oriented ML artifact. Results render into the `## Validation` section of `project-specs.md` per `shared/validation_protocol.md`.
4
+
5
+ Check IDs (AMS-01 through AMS-09) are stable. Applied ML Science validation is closer to empirical research than to production engineering: the emphasis is on soundness of claims, reproducibility of results, and rigorous baselines — not deployment readiness.
6
+
7
+ ## AMS-01 — Research Question Precisely Stated
8
+
9
+ The question being answered is specific, falsifiable, and scoped.
10
+
11
+ - Statement includes: the phenomenon studied, the hypothesis being tested, the metric that would confirm or refute it, and the scope boundary.
12
+ - Vague goals ("improve performance") are rewritten as specific ("reduce validation loss by ≥5% vs baseline on benchmark X under compute budget Y").
13
+
14
+ **Observed format:** `question: "Does contrastive pretraining on task-adjacent unlabeled data reduce few-shot classification error on benchmark B by ≥10% at 100 labels, relative to supervised-only baseline?" | scope: benchmark B only; claims do not generalize to other benchmarks without separate validation`
15
+
16
+ ## AMS-02 — Baselines Rigorous
17
+
18
+ Comparisons include both trivial baselines and strong SOTA baselines where applicable.
19
+
20
+ - Trivial: random, majority class, nearest-neighbor, linear probe.
21
+ - Strong: current best-known published method or the best internal method — reproduced in the same evaluation harness, not quoted from paper.
22
+ - Baseline code and configs on disk; runs repeatable.
23
+
24
+ **Observed format:** `baselines: linear probe (trivial), SimCLR v2 (SOTA reproduced locally on same harness), supervised-only (direct comparison) | baseline configs: configs/baselines/ | reported: all three on identical eval splits`
25
+
26
+ ## AMS-03 — Ablation Studies
27
+
28
+ The novel method's claimed source of improvement is isolated.
29
+
30
+ - For every non-trivial design choice claimed to contribute (loss term, architecture component, training technique): an ablation run with that choice removed/altered.
31
+ - Each ablation reports the headline metric, so contribution can be attributed.
32
+ - Negative-result ablations (design choices that didn't help) also reported.
33
+
34
+ **Observed format:** `5 design choices → 5 ablations run | contribution breakdown: contrastive_loss +4.2pp, temperature=0.1 +1.8pp, projection_head +0.7pp, aug_policy +2.1pp, batch_size ≥1024 +0.5pp | negative: momentum_encoder -0.3pp (removed from final) | results/ablations.json`
35
+
36
+ ## AMS-04 — Statistical Significance of Improvements
37
+
38
+ Claimed improvements are statistically meaningful given the variance of the training setup.
39
+
40
+ - Multiple training runs with different seeds (≥3, ideally ≥5 for small-data regimes).
41
+ - Mean ± std reported, not single-run numbers.
42
+ - Paired tests or bootstrapped CIs used when comparing methods — a single percentage-point improvement within run-to-run variance is not a finding.
43
+
44
+ **Observed format:** `5 seeds per method | novel: 78.4% ± 0.8% | SOTA baseline: 74.1% ± 1.1% | paired t-test p=0.003, 95% CI on delta = [3.1%, 5.5%] | results/seed_variance.json`
45
+
46
+ ## AMS-05 — Theoretical Soundness (for novel methods)
47
+
48
+ For frameworks with theoretical claims, the math checks out and assumptions are stated.
49
+
50
+ - Derivations reviewed for errors. Consult the Researcher via Task for statistical methodology claims.
51
+ - Assumptions named (stationarity, independence, convexity, smoothness, i.i.d.).
52
+ - Counterexamples to claimed properties probed where feasible.
53
+ - Empirical results consistent with theoretical predictions (or discrepancy explained).
54
+
55
+ **Observed format:** `derivation: notes/derivation.md §2-4, reviewed by Researcher via Task (verdict APPROVED) | assumptions: data is i.i.d. sub-Gaussian (stated), bounded loss, smooth encoder | empirical-theoretical gap: convergence rate ~O(1/√T) matches theory within constants ✓`
56
+
57
+ Skip with `n/a` for purely empirical studies with no theoretical claims.
58
+
59
+ ## AMS-06 — Reproducibility: Seed, Environment, Data
60
+
61
+ A collaborator (or future-you in three months) can reproduce the headline numbers from the committed artifacts.
62
+
63
+ - All seeds pinned (data split, model init, trainer, any augmentation sampler).
64
+ - Environment captured: `requirements.txt` with pinned versions, hardware spec, CUDA version where relevant.
65
+ - Data: exact split definition on disk (manifest of IDs or deterministic split rule).
66
+ - A single `README.md` command reproduces the headline number.
67
+
68
+ **Observed format:** `seeds=[42,43,44,45,46] | env: research/<project>/env/requirements.txt (pinned) | hardware: 4×A100 80GB, CUDA 12.1 | data manifest: data/splits/manifest_v3.json | repro command: make repro in README; re-run produced 78.3% vs reported 78.4% (within seed variance) ✓`
69
+
70
+ ## AMS-07 — Training Dynamics Documented
71
+
72
+ Training curves, gradient behavior, and any instabilities are recorded — not just the final number.
73
+
74
+ - Loss curves (train + val), gradient norms, learning rate schedule captured in logs.
75
+ - Instabilities (NaNs, divergence, plateaus) noted and their treatment described.
76
+ - If training was unstable to reproduce (large seed variance), that is itself the finding — do not hide it.
77
+
78
+ **Observed format:** `W&B run IDs: [run_a8f, run_b2c, run_3dd, run_91e, run_44f] | loss curves monotone after epoch 5, val plateau epoch 80 (used for early stop) | 1 NaN observed on seed=44 run (gradient clip threshold too loose initially, fixed) | plots: results/training_curves.png`
79
+
80
+ ## AMS-08 — Code on Disk + Component Tests
81
+
82
+ The research code is written as testable modules, not notebook-only, and key components have tests.
83
+
84
+ - Research code under version control (or at least a clear source-of-truth directory).
85
+ - Unit tests for: loss functions, custom modules, data augmentation, eval scorers.
86
+ - "It ran once in my notebook" is not enough — research that cannot be re-run in a clean environment is not validated.
87
+
88
+ **Observed format:** `research/<project>/src/ on disk (not notebook-only) | tests/: 17 tests, 17 passed | covered: contrastive_loss (forward + gradient), projection_head (output shape + init), aug_pipeline (determinism with seed), eval_scorer (parity with reference impl)`
89
+
90
+ ## AMS-09 — Scope and Negative Claims
91
+
92
+ Scope boundaries are stated, and negative results or known failure modes are reported.
93
+
94
+ - Where does the method work? Where has it been tested?
95
+ - Where does it not work, or where is it untested? (Different domain, different scale, different task.)
96
+ - Any negative results from the study are reported, not suppressed.
97
+
98
+ **Observed format:** `scope: benchmark B, scale 10k-100k labels, image-classification task family | untested: language, tabular, outside-scale | negative result: on fine-grained subset (benchmark B-fine), method regresses by 2.3pp — reported in paper §6 | limits discussion: report §7`
99
+
100
+ ---
101
+
102
+ ## Track Calibration
103
+
104
+ Rows are indexed by `(Track, Mode)` per `shared/validation_protocol.md`.
105
+
106
+ | Track | Mode | Required | Recommended | Skippable |
107
+ |-------|------|----------|-------------|-----------|
108
+ | **deep** | `create` (novel framework from scratch) | AMS-01, AMS-02, AMS-03, AMS-04, AMS-06, AMS-07, AMS-08, AMS-09 | AMS-05 | — |
109
+ | **deep** | `review` (methodology review / advisory writeup) | AMS-01, AMS-02, AMS-05, AMS-09 | AMS-03 | AMS-04, AMS-06, AMS-07, AMS-08 (no new artifacts) |
110
+ | **quick** | `experiment` (kept `[X]` iteration) | AMS-04 (seed variance for the kept change) + diff vs prior | AMS-07 | most |
111
+ | **fixer** | (Mode omitted) | AMS-08 + "what changed, what didn't break" | — | rest |
112
+
113
+ Any skipped or inapplicable check must still appear as a row with `Pass/Fail: n/a` and a Notes cell giving the reason. See `shared/validation_protocol.md`.
114
+
115
+ ## Artifacts Expected
116
+
117
+ - Research code directory under `research/<project>/` — AMS-08
118
+ - `tests/` directory — AMS-08
119
+ - `results/ablations.json`, `results/seed_variance.json` — AMS-03, AMS-04
120
+ - `results/training_curves.png` + W&B/MLflow run manifest — AMS-07
121
+ - `README.md` with reproduction command — AMS-06
122
+ - Paper/report draft referencing all evidence — consolidates the claims
123
+
124
+ ## Downstream Impact — What to Cover
125
+
126
+ - **Production adopters:** if the method will be productionized, flag for ML Engineer and Deep Learning Engineer — this checklist is research-grade, not production-grade.
127
+ - **Published claims:** if results will be published, Academic review via Task before release.
128
+ - **Shared infrastructure:** if the method imposes new compute requirements, flag for MLOps.
129
+
130
+ ## When to Escalate
131
+
132
+ - **AMS-04 improvement falls within seed variance** — claim is not supported; do not ship as a positive result. Either run more seeds or reframe as exploratory.
133
+ - **AMS-05 theoretical claims fail review** — rewrite as empirical with no theoretical framing, or retract the claim.
134
+ - **AMS-06 reproduction fails (headline number can't be reproduced)** — halt. This is the most serious failure mode in research; find and fix the source of non-determinism before making any claims.
135
+ - **AMS-09 scope claims that can't be defended** — narrow the scope until they can.
136
+ - **Any check produces a result the agent cannot explain.** Record as `✗` and surface in Open Issues.