@proflandrigan/shards 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (397) hide show
  1. package/README.md +475 -0
  2. package/package.json +37 -0
  3. package/src/agents/academic.md +276 -0
  4. package/src/agents/ai-engineer.md +377 -0
  5. package/src/agents/analytics-engineer.md +364 -0
  6. package/src/agents/applied-ml-scientist.md +410 -0
  7. package/src/agents/backend-engineer.md +255 -0
  8. package/src/agents/bi-engineer.md +333 -0
  9. package/src/agents/data-analyst.md +343 -0
  10. package/src/agents/data-engineer.md +260 -0
  11. package/src/agents/data-modeller.md +386 -0
  12. package/src/agents/data-scientist.md +366 -0
  13. package/src/agents/deep-learning-engineer.md +389 -0
  14. package/src/agents/ml-engineer.md +424 -0
  15. package/src/agents/mlops-engineer.md +339 -0
  16. package/src/agents/researcher.md +187 -0
  17. package/src/agents/specific_instructions/academic/critical_review.md +263 -0
  18. package/src/agents/specific_instructions/academic/report.md +113 -0
  19. package/src/agents/specific_instructions/ai_engineer/advise.md +162 -0
  20. package/src/agents/specific_instructions/ai_engineer/bi_engineer_handoff.md +86 -0
  21. package/src/agents/specific_instructions/ai_engineer/experiment.md +471 -0
  22. package/src/agents/specific_instructions/ai_engineer/experiment_ui_mode.md +44 -0
  23. package/src/agents/specific_instructions/ai_engineer/phases/index.md +45 -0
  24. package/src/agents/specific_instructions/ai_engineer/phases/phase-1.md +55 -0
  25. package/src/agents/specific_instructions/ai_engineer/phases/phase-2.md +86 -0
  26. package/src/agents/specific_instructions/ai_engineer/phases/phase-3.md +96 -0
  27. package/src/agents/specific_instructions/ai_engineer/phases/phase-4.md +138 -0
  28. package/src/agents/specific_instructions/ai_engineer/phases/phase-5.md +157 -0
  29. package/src/agents/specific_instructions/ai_engineer/phases/phase-6.md +196 -0
  30. package/src/agents/specific_instructions/ai_engineer/phases/phase-7.md +313 -0
  31. package/src/agents/specific_instructions/ai_engineer/phases.md +1011 -0
  32. package/src/agents/specific_instructions/ai_engineer/prompt_lab.md +161 -0
  33. package/src/agents/specific_instructions/ai_engineer/prompt_lab_ui_mode.md +28 -0
  34. package/src/agents/specific_instructions/ai_engineer/research.md +393 -0
  35. package/src/agents/specific_instructions/ai_engineer/research_ui_mode.md +66 -0
  36. package/src/agents/specific_instructions/ai_engineer/review.md +159 -0
  37. package/src/agents/specific_instructions/ai_engineer/validation_checklist.md +182 -0
  38. package/src/agents/specific_instructions/analytics_engineer/advise.md +155 -0
  39. package/src/agents/specific_instructions/analytics_engineer/bi_engineer_handoff.md +91 -0
  40. package/src/agents/specific_instructions/analytics_engineer/data_analyst_handoff.md +84 -0
  41. package/src/agents/specific_instructions/analytics_engineer/deep_phases.md +818 -0
  42. package/src/agents/specific_instructions/analytics_engineer/phases_deep/index.md +24 -0
  43. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-1.md +77 -0
  44. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-2.md +106 -0
  45. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-3.md +93 -0
  46. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-4.md +79 -0
  47. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-5.md +61 -0
  48. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-6.md +45 -0
  49. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-7.md +235 -0
  50. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-8.md +221 -0
  51. package/src/agents/specific_instructions/analytics_engineer/phases_quick/index.md +19 -0
  52. package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-1.md +47 -0
  53. package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-2.md +78 -0
  54. package/src/agents/specific_instructions/analytics_engineer/quick_phases.md +112 -0
  55. package/src/agents/specific_instructions/analytics_engineer/review.md +167 -0
  56. package/src/agents/specific_instructions/analytics_engineer/service_mode.md +369 -0
  57. package/src/agents/specific_instructions/analytics_engineer/ui_mode.md +45 -0
  58. package/src/agents/specific_instructions/analytics_engineer/update.md +162 -0
  59. package/src/agents/specific_instructions/analytics_engineer/validation_checklist.md +121 -0
  60. package/src/agents/specific_instructions/applied_ml_scientist/advise.md +143 -0
  61. package/src/agents/specific_instructions/applied_ml_scientist/phases/index.md +21 -0
  62. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-1.md +51 -0
  63. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-2.md +66 -0
  64. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-3.md +113 -0
  65. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-4.md +104 -0
  66. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-5.md +156 -0
  67. package/src/agents/specific_instructions/applied_ml_scientist/phases.md +428 -0
  68. package/src/agents/specific_instructions/applied_ml_scientist/research.md +379 -0
  69. package/src/agents/specific_instructions/applied_ml_scientist/review.md +142 -0
  70. package/src/agents/specific_instructions/applied_ml_scientist/validation_checklist.md +136 -0
  71. package/src/agents/specific_instructions/backend_engineer/clean.md +149 -0
  72. package/src/agents/specific_instructions/backend_engineer/review.md +91 -0
  73. package/src/agents/specific_instructions/backend_engineer/review_checklist.md +54 -0
  74. package/src/agents/specific_instructions/backend_engineer/service_mode.md +67 -0
  75. package/src/agents/specific_instructions/bi_engineer/advise.md +137 -0
  76. package/src/agents/specific_instructions/bi_engineer/data_analyst_handoff.md +77 -0
  77. package/src/agents/specific_instructions/bi_engineer/incoming_handoff.md +45 -0
  78. package/src/agents/specific_instructions/bi_engineer/phases/index.md +20 -0
  79. package/src/agents/specific_instructions/bi_engineer/phases/phase-1.md +164 -0
  80. package/src/agents/specific_instructions/bi_engineer/phases/phase-2.md +92 -0
  81. package/src/agents/specific_instructions/bi_engineer/phases/phase-3.md +121 -0
  82. package/src/agents/specific_instructions/bi_engineer/phases/phase-4.md +106 -0
  83. package/src/agents/specific_instructions/bi_engineer/phases.md +451 -0
  84. package/src/agents/specific_instructions/bi_engineer/review.md +166 -0
  85. package/src/agents/specific_instructions/bi_engineer/update.md +147 -0
  86. package/src/agents/specific_instructions/bi_engineer/validation_checklist.md +124 -0
  87. package/src/agents/specific_instructions/data_analyst/advise.md +138 -0
  88. package/src/agents/specific_instructions/data_analyst/explain.md +221 -0
  89. package/src/agents/specific_instructions/data_analyst/incoming_handoff.md +40 -0
  90. package/src/agents/specific_instructions/data_analyst/phases/index.md +20 -0
  91. package/src/agents/specific_instructions/data_analyst/phases/phase-1.md +159 -0
  92. package/src/agents/specific_instructions/data_analyst/phases/phase-2.md +112 -0
  93. package/src/agents/specific_instructions/data_analyst/phases/phase-3.md +265 -0
  94. package/src/agents/specific_instructions/data_analyst/phases/phase-4.md +100 -0
  95. package/src/agents/specific_instructions/data_analyst/phases.md +501 -0
  96. package/src/agents/specific_instructions/data_analyst/review.md +138 -0
  97. package/src/agents/specific_instructions/data_analyst/ui_mode.md +26 -0
  98. package/src/agents/specific_instructions/data_analyst/update.md +144 -0
  99. package/src/agents/specific_instructions/data_analyst/validation_checklist.md +95 -0
  100. package/src/agents/specific_instructions/data_engineer/advise.md +137 -0
  101. package/src/agents/specific_instructions/data_engineer/phases.md +466 -0
  102. package/src/agents/specific_instructions/data_engineer/phases_deep/index.md +23 -0
  103. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-1.md +49 -0
  104. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-2.md +93 -0
  105. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-3.md +55 -0
  106. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-4.md +48 -0
  107. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-5.md +40 -0
  108. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-6.md +102 -0
  109. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-7.md +87 -0
  110. package/src/agents/specific_instructions/data_engineer/phases_quick/index.md +19 -0
  111. package/src/agents/specific_instructions/data_engineer/phases_quick/phase-1.md +45 -0
  112. package/src/agents/specific_instructions/data_engineer/phases_quick/phase-2.md +54 -0
  113. package/src/agents/specific_instructions/data_engineer/review.md +135 -0
  114. package/src/agents/specific_instructions/data_engineer/validation_checklist.md +136 -0
  115. package/src/agents/specific_instructions/data_modeller/advise.md +137 -0
  116. package/src/agents/specific_instructions/data_modeller/phases.md +581 -0
  117. package/src/agents/specific_instructions/data_modeller/phases_deep/index.md +23 -0
  118. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-1.md +52 -0
  119. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-2.md +113 -0
  120. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-3.md +47 -0
  121. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-4.md +51 -0
  122. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-5.md +45 -0
  123. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-6.md +105 -0
  124. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-7.md +136 -0
  125. package/src/agents/specific_instructions/data_modeller/phases_quick/index.md +19 -0
  126. package/src/agents/specific_instructions/data_modeller/phases_quick/phase-1.md +47 -0
  127. package/src/agents/specific_instructions/data_modeller/phases_quick/phase-2.md +65 -0
  128. package/src/agents/specific_instructions/data_modeller/review.md +141 -0
  129. package/src/agents/specific_instructions/data_modeller/service_mode.md +218 -0
  130. package/src/agents/specific_instructions/data_modeller/validation_checklist.md +125 -0
  131. package/src/agents/specific_instructions/data_scientist/advise.md +158 -0
  132. package/src/agents/specific_instructions/data_scientist/bi_engineer_handoff.md +63 -0
  133. package/src/agents/specific_instructions/data_scientist/experiment.md +482 -0
  134. package/src/agents/specific_instructions/data_scientist/experiment_ui_mode.md +44 -0
  135. package/src/agents/specific_instructions/data_scientist/explain.md +247 -0
  136. package/src/agents/specific_instructions/data_scientist/greenfield_data.md +35 -0
  137. package/src/agents/specific_instructions/data_scientist/ml_engineer_handoff.md +52 -0
  138. package/src/agents/specific_instructions/data_scientist/notebook_walkthrough.md +76 -0
  139. package/src/agents/specific_instructions/data_scientist/phases/index.md +24 -0
  140. package/src/agents/specific_instructions/data_scientist/phases/phase-1.md +45 -0
  141. package/src/agents/specific_instructions/data_scientist/phases/phase-2.md +67 -0
  142. package/src/agents/specific_instructions/data_scientist/phases/phase-3.md +89 -0
  143. package/src/agents/specific_instructions/data_scientist/phases/phase-4.md +143 -0
  144. package/src/agents/specific_instructions/data_scientist/phases/phase-5.md +71 -0
  145. package/src/agents/specific_instructions/data_scientist/phases/phase-6.md +239 -0
  146. package/src/agents/specific_instructions/data_scientist/phases/phase-7.md +207 -0
  147. package/src/agents/specific_instructions/data_scientist/phases.md +651 -0
  148. package/src/agents/specific_instructions/data_scientist/research.md +345 -0
  149. package/src/agents/specific_instructions/data_scientist/research_ui_mode.md +52 -0
  150. package/src/agents/specific_instructions/data_scientist/review.md +136 -0
  151. package/src/agents/specific_instructions/data_scientist/service_mode.md +247 -0
  152. package/src/agents/specific_instructions/data_scientist/validation_checklist.md +183 -0
  153. package/src/agents/specific_instructions/deep_learning_engineer/advise.md +145 -0
  154. package/src/agents/specific_instructions/deep_learning_engineer/phases/index.md +21 -0
  155. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-1.md +74 -0
  156. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-2.md +98 -0
  157. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-3.md +76 -0
  158. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-4.md +128 -0
  159. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-5.md +292 -0
  160. package/src/agents/specific_instructions/deep_learning_engineer/phases.md +567 -0
  161. package/src/agents/specific_instructions/deep_learning_engineer/research.md +389 -0
  162. package/src/agents/specific_instructions/deep_learning_engineer/review.md +155 -0
  163. package/src/agents/specific_instructions/deep_learning_engineer/validation_checklist.md +147 -0
  164. package/src/agents/specific_instructions/ml_engineer/advise.md +174 -0
  165. package/src/agents/specific_instructions/ml_engineer/bi_engineer_handoff.md +71 -0
  166. package/src/agents/specific_instructions/ml_engineer/experiment.md +474 -0
  167. package/src/agents/specific_instructions/ml_engineer/experiment_ui_mode.md +44 -0
  168. package/src/agents/specific_instructions/ml_engineer/notebook_walkthrough.md +75 -0
  169. package/src/agents/specific_instructions/ml_engineer/phases/index.md +25 -0
  170. package/src/agents/specific_instructions/ml_engineer/phases/phase-1.md +49 -0
  171. package/src/agents/specific_instructions/ml_engineer/phases/phase-2.md +75 -0
  172. package/src/agents/specific_instructions/ml_engineer/phases/phase-3.md +124 -0
  173. package/src/agents/specific_instructions/ml_engineer/phases/phase-4.md +279 -0
  174. package/src/agents/specific_instructions/ml_engineer/phases/phase-5.md +160 -0
  175. package/src/agents/specific_instructions/ml_engineer/phases/phase-6-5.md +170 -0
  176. package/src/agents/specific_instructions/ml_engineer/phases/phase-6.md +295 -0
  177. package/src/agents/specific_instructions/ml_engineer/phases/phase-7.md +337 -0
  178. package/src/agents/specific_instructions/ml_engineer/phases.md +1068 -0
  179. package/src/agents/specific_instructions/ml_engineer/research.md +437 -0
  180. package/src/agents/specific_instructions/ml_engineer/research_ui_mode.md +71 -0
  181. package/src/agents/specific_instructions/ml_engineer/review.md +187 -0
  182. package/src/agents/specific_instructions/ml_engineer/service_mode.md +273 -0
  183. package/src/agents/specific_instructions/ml_engineer/validation_checklist.md +185 -0
  184. package/src/agents/specific_instructions/mlops_engineer/advise.md +139 -0
  185. package/src/agents/specific_instructions/mlops_engineer/phases/index.md +23 -0
  186. package/src/agents/specific_instructions/mlops_engineer/phases/phase-1.md +52 -0
  187. package/src/agents/specific_instructions/mlops_engineer/phases/phase-2.md +86 -0
  188. package/src/agents/specific_instructions/mlops_engineer/phases/phase-3.md +105 -0
  189. package/src/agents/specific_instructions/mlops_engineer/phases/phase-4.md +128 -0
  190. package/src/agents/specific_instructions/mlops_engineer/phases/phase-5.md +106 -0
  191. package/src/agents/specific_instructions/mlops_engineer/phases/phase-6.md +128 -0
  192. package/src/agents/specific_instructions/mlops_engineer/phases/phase-7.md +144 -0
  193. package/src/agents/specific_instructions/mlops_engineer/phases.md +671 -0
  194. package/src/agents/specific_instructions/mlops_engineer/review.md +164 -0
  195. package/src/agents/specific_instructions/mlops_engineer/service_mode.md +81 -0
  196. package/src/agents/specific_instructions/mlops_engineer/validation_checklist.md +151 -0
  197. package/src/agents/specific_instructions/researcher/critical_review.md +292 -0
  198. package/src/agents/specific_instructions/researcher/review_checklist.md +67 -0
  199. package/src/agents/specific_instructions/researcher/service_mode.md +224 -0
  200. package/src/agents/specific_instructions/shared/auto_verify_mode.md +141 -0
  201. package/src/agents/specific_instructions/shared/autonomous_research.md +1289 -0
  202. package/src/agents/specific_instructions/shared/behavioral_rules.md +36 -0
  203. package/src/agents/specific_instructions/shared/diverge_protocol.md +387 -0
  204. package/src/agents/specific_instructions/shared/engineering_guidelines.md +136 -0
  205. package/src/agents/specific_instructions/shared/experiment_versioning.md +184 -0
  206. package/src/agents/specific_instructions/shared/goal_mode.md +187 -0
  207. package/src/agents/specific_instructions/shared/incremental_testing.md +139 -0
  208. package/src/agents/specific_instructions/shared/intent_discovery.md +223 -0
  209. package/src/agents/specific_instructions/shared/join_path_protocol.md +168 -0
  210. package/src/agents/specific_instructions/shared/knowledge_checkpoint.md +83 -0
  211. package/src/agents/specific_instructions/shared/knowledge_harvest.md +220 -0
  212. package/src/agents/specific_instructions/shared/knowledge_retrieval.md +100 -0
  213. package/src/agents/specific_instructions/shared/notebook_walkthrough_protocol.md +367 -0
  214. package/src/agents/specific_instructions/shared/reviewer_verdict_protocol.md +74 -0
  215. package/src/agents/specific_instructions/shared/swarm_protocol.md +97 -0
  216. package/src/agents/specific_instructions/shared/validation_protocol.md +139 -0
  217. package/src/agents/specific_instructions/syn/arbiter.md +140 -0
  218. package/src/agents/specific_instructions/syn/brainstorm.md +550 -0
  219. package/src/agents/specific_instructions/syn/code_review.md +232 -0
  220. package/src/agents/specific_instructions/syn/diff.md +239 -0
  221. package/src/agents/specific_instructions/syn/final_review.md +65 -0
  222. package/src/agents/specific_instructions/syn/fixer.md +240 -0
  223. package/src/agents/specific_instructions/syn/free_form.md +130 -0
  224. package/src/agents/specific_instructions/syn/knowledge.md +468 -0
  225. package/src/agents/specific_instructions/syn/notebook_walkthrough.md +78 -0
  226. package/src/agents/specific_instructions/syn/panel_review.md +634 -0
  227. package/src/agents/specific_instructions/syn/pm.md +453 -0
  228. package/src/agents/specific_instructions/syn/pr_review.md +255 -0
  229. package/src/agents/specific_instructions/syn/slides.md +417 -0
  230. package/src/agents/syn.md +729 -0
  231. package/src/commands/academic.md +41 -0
  232. package/src/commands/ai-engineer.md +45 -0
  233. package/src/commands/analytics-engineer.md +48 -0
  234. package/src/commands/applied-ml-scientist.md +45 -0
  235. package/src/commands/backend-engineer.md +35 -0
  236. package/src/commands/bi-engineer.md +40 -0
  237. package/src/commands/brainstorm.md +24 -0
  238. package/src/commands/data-analyst.md +38 -0
  239. package/src/commands/data-engineer.md +37 -0
  240. package/src/commands/data-modeller.md +38 -0
  241. package/src/commands/data-scientist.md +38 -0
  242. package/src/commands/deep-learning-engineer.md +47 -0
  243. package/src/commands/end.md +49 -0
  244. package/src/commands/knowledge.md +24 -0
  245. package/src/commands/ml-engineer.md +42 -0
  246. package/src/commands/mlops-engineer.md +47 -0
  247. package/src/commands/notebook-walkthrough.md +58 -0
  248. package/src/commands/researcher.md +40 -0
  249. package/src/commands/resume.md +57 -0
  250. package/src/commands/review-pr.md +26 -0
  251. package/src/commands/shards-guide.md +41 -0
  252. package/src/commands/shards-ui.md +32 -0
  253. package/src/commands/shards.md +41 -0
  254. package/src/docs/01-getting-started/concepts.md +109 -0
  255. package/src/docs/01-getting-started/first-session.md +79 -0
  256. package/src/docs/01-getting-started/install.md +61 -0
  257. package/src/docs/02-agents/academic.md +71 -0
  258. package/src/docs/02-agents/ai-engineer.md +78 -0
  259. package/src/docs/02-agents/analytics-engineer.md +58 -0
  260. package/src/docs/02-agents/applied-ml-scientist.md +59 -0
  261. package/src/docs/02-agents/backend-engineer.md +58 -0
  262. package/src/docs/02-agents/bi-engineer.md +65 -0
  263. package/src/docs/02-agents/data-analyst.md +67 -0
  264. package/src/docs/02-agents/data-engineer.md +57 -0
  265. package/src/docs/02-agents/data-modeller.md +51 -0
  266. package/src/docs/02-agents/data-scientist.md +78 -0
  267. package/src/docs/02-agents/deep-learning-engineer.md +64 -0
  268. package/src/docs/02-agents/ml-engineer.md +80 -0
  269. package/src/docs/02-agents/mlops-engineer.md +59 -0
  270. package/src/docs/02-agents/overview.md +62 -0
  271. package/src/docs/02-agents/researcher.md +73 -0
  272. package/src/docs/02-agents/syn.md +88 -0
  273. package/src/docs/03-protocols/auto-verify.md +82 -0
  274. package/src/docs/03-protocols/autonomous-research.md +59 -0
  275. package/src/docs/03-protocols/behavioral-rules.md +35 -0
  276. package/src/docs/03-protocols/diverge.md +50 -0
  277. package/src/docs/03-protocols/engineering-guidelines.md +56 -0
  278. package/src/docs/03-protocols/experiment-versioning.md +38 -0
  279. package/src/docs/03-protocols/gate-pattern.md +65 -0
  280. package/src/docs/03-protocols/incremental-testing.md +68 -0
  281. package/src/docs/03-protocols/join-path.md +46 -0
  282. package/src/docs/03-protocols/knowledge-ledger.md +70 -0
  283. package/src/docs/03-protocols/reviewer-verdicts.md +39 -0
  284. package/src/docs/03-protocols/swarm.md +40 -0
  285. package/src/docs/03-protocols/validation.md +174 -0
  286. package/src/docs/04-ui/activity-bar.md +70 -0
  287. package/src/docs/04-ui/chat-pane.md +80 -0
  288. package/src/docs/04-ui/code-intel.md +62 -0
  289. package/src/docs/04-ui/file-editing.md +61 -0
  290. package/src/docs/04-ui/git.md +54 -0
  291. package/src/docs/04-ui/keybindings.md +79 -0
  292. package/src/docs/04-ui/knowledge-map.md +76 -0
  293. package/src/docs/04-ui/overview.md +93 -0
  294. package/src/docs/04-ui/panels.md +49 -0
  295. package/src/docs/04-ui/pinboard-selection.md +66 -0
  296. package/src/docs/04-ui/quick-open-palette.md +56 -0
  297. package/src/docs/04-ui/sessions.md +81 -0
  298. package/src/docs/04-ui/settings-permissions.md +56 -0
  299. package/src/docs/05-commands/reference.md +59 -0
  300. package/src/docs/06-outputs/directory-map.md +116 -0
  301. package/src/docs/07-workflows/ai-eval-first.md +57 -0
  302. package/src/docs/07-workflows/deep-study-to-production.md +76 -0
  303. package/src/docs/07-workflows/diverge-exploration.md +77 -0
  304. package/src/docs/07-workflows/quick-analysis.md +45 -0
  305. package/src/docs/08-integrations/claude-code-auto-mode.md +191 -0
  306. package/src/docs/08-integrations/google-slides.md +175 -0
  307. package/src/docs/README.md +30 -0
  308. package/src/docs/manifest.json +108 -0
  309. package/src/templates/analysis-template.md +20 -0
  310. package/src/templates/branch-report.md +46 -0
  311. package/src/templates/diff-report.md +88 -0
  312. package/src/templates/knowledge-index.md +7 -0
  313. package/src/templates/model-card-schema.json +186 -0
  314. package/src/templates/model-card-schema.md +88 -0
  315. package/src/templates/model-card.md +124 -0
  316. package/src/templates/project-plan.md +47 -0
  317. package/src/templates/project-specs.md +81 -0
  318. package/src/templates/report-template.md +43 -0
  319. package/src/templates/study-template.md +25 -0
  320. package/src/ui/cc-readonly.js +181 -0
  321. package/src/ui/chat-session.js +466 -0
  322. package/src/ui/css/base.css +136 -0
  323. package/src/ui/css/brainstorm.css +525 -0
  324. package/src/ui/css/chat.css +1405 -0
  325. package/src/ui/css/editor.css +546 -0
  326. package/src/ui/css/eval-dashboard.css +157 -0
  327. package/src/ui/css/experiment.css +237 -0
  328. package/src/ui/css/guide.css +186 -0
  329. package/src/ui/css/knowledge-map.css +383 -0
  330. package/src/ui/css/layout.css +431 -0
  331. package/src/ui/css/model-card.css +161 -0
  332. package/src/ui/css/notebook-walkthrough.css +271 -0
  333. package/src/ui/css/pr-review.css +403 -0
  334. package/src/ui/css/prompt-lab.css +325 -0
  335. package/src/ui/css/sessions.css +258 -0
  336. package/src/ui/css/sidebar.css +661 -0
  337. package/src/ui/css/terminal.css +113 -0
  338. package/src/ui/css/theme-light.css +542 -0
  339. package/src/ui/index.html +389 -0
  340. package/src/ui/js/agents.js +32 -0
  341. package/src/ui/js/bookmarks.js +230 -0
  342. package/src/ui/js/chat.js +1776 -0
  343. package/src/ui/js/code-intel.js +328 -0
  344. package/src/ui/js/command-palette.js +142 -0
  345. package/src/ui/js/events.js +591 -0
  346. package/src/ui/js/explorer.js +317 -0
  347. package/src/ui/js/file-view.js +477 -0
  348. package/src/ui/js/git.js +536 -0
  349. package/src/ui/js/guide.js +198 -0
  350. package/src/ui/js/hud.js +75 -0
  351. package/src/ui/js/init.js +351 -0
  352. package/src/ui/js/knowledge-map.js +906 -0
  353. package/src/ui/js/markdown.js +114 -0
  354. package/src/ui/js/monaco.js +164 -0
  355. package/src/ui/js/notebook-walkthrough.js +272 -0
  356. package/src/ui/js/notebook.js +448 -0
  357. package/src/ui/js/panels.js +2681 -0
  358. package/src/ui/js/pinboard.js +186 -0
  359. package/src/ui/js/quick-open.js +164 -0
  360. package/src/ui/js/selection-context.js +131 -0
  361. package/src/ui/js/sessions.js +256 -0
  362. package/src/ui/js/settings.js +476 -0
  363. package/src/ui/js/split-view.js +82 -0
  364. package/src/ui/js/state.js +343 -0
  365. package/src/ui/js/table.js +161 -0
  366. package/src/ui/js/tabs.js +284 -0
  367. package/src/ui/js/tabular.js +125 -0
  368. package/src/ui/js/terminal.js +354 -0
  369. package/src/ui/js/timeline.js +137 -0
  370. package/src/ui/js/utils.js +293 -0
  371. package/src/ui/notebook-kernel.py +790 -0
  372. package/src/ui/open-browser.js +55 -0
  373. package/src/ui/permission-pattern.js +42 -0
  374. package/src/ui/relay.js +513 -0
  375. package/src/ui/server.js +3072 -0
  376. package/src/ui/session-index.js +225 -0
  377. package/src/ui/shards_icon.png +0 -0
  378. package/src/ui/spawn-server.js +41 -0
  379. package/src/ui/symbol-index.js +813 -0
  380. package/src/ui/ui-push.js +177 -0
  381. package/tools/gate-hook/VALIDATION_SPEC.md +273 -0
  382. package/tools/gate-hook/__tests__/auto-verify.test.js +343 -0
  383. package/tools/gate-hook/auto-allowlist.js +179 -0
  384. package/tools/gate-hook/auto-state.js +68 -0
  385. package/tools/gate-hook/classify.js +21 -0
  386. package/tools/gate-hook/log.js +57 -0
  387. package/tools/gate-hook/parser.js +205 -0
  388. package/tools/gate-hook/sql-guard.js +230 -0
  389. package/tools/gate-hook/state.js +170 -0
  390. package/tools/gate-hook/sweep.js +139 -0
  391. package/tools/gate-hook/transcript.js +45 -0
  392. package/tools/gate-hook/validation.js +321 -0
  393. package/tools/gate-hook.js +475 -0
  394. package/tools/install.js +914 -0
  395. package/tools/shards-gates.js +311 -0
  396. package/tools/shards-sessions.js +261 -0
  397. package/tools/shards-ui.js +377 -0
@@ -0,0 +1,86 @@
1
+ > **Previous:** phase-1.md confirmed
2
+ > **Next:** phase-3.md (read only after this phase's gate is confirmed)
3
+
4
+ ---
5
+
6
+ ## Phase 2 — Infrastructure Assessment
7
+
8
+ Goal: Understand existing infrastructure and constraints before designing anything.
9
+
10
+ Ask about:
11
+ - **Existing ML infrastructure:** model registry, feature store, serving layer,
12
+ training orchestrator, experiment tracker — what exists, what doesn't
13
+ - **Cloud services available:** Which managed services can we use? Any
14
+ organizational restrictions?
15
+ - **Compliance / security constraints:** Data residency requirements, VPC
16
+ isolation, IAM constraints, PII handling, audit logging requirements
17
+ - **Team capabilities:** What tooling does the team already know and operate?
18
+ (Important — the right tool for a team that knows Kubernetes is different from
19
+ the right tool for a team that doesn't)
20
+ - **Data pipeline integration:** How do features arrive at training time?
21
+ At serving time? What's the existing data infrastructure?
22
+ - **Existing monitoring:** Any existing observability stack (Prometheus,
23
+ Datadog, CloudWatch, etc.) that ML monitoring should integrate with?
24
+
25
+ **Consult the ML Engineer** for model architecture constraints that affect serving:
26
+
27
+ Tell the user: "Getting the ML Engineer in here — I need to know what the model actually requires before I design serving infrastructure around assumptions."
28
+
29
+ ```
30
+ Task(
31
+ subagent_type="ml-engineer",
32
+ description="Review model architecture constraints for MLOps serving design",
33
+ prompt="I am the MLOps Engineer shard scoping an MLOps project for: [project description].
34
+ I need to understand the model architecture constraints that affect my serving and
35
+ infrastructure design. Please tell me:
36
+ 1. What is the model framework and format (scikit-learn, XGBoost, PyTorch, TensorFlow, etc.)?
37
+ 2. What are the model size and memory requirements (serialized size, memory at inference)?
38
+ 3. What are the serving-time feature requirements (features needed at inference, latency sensitivity)?
39
+ 4. Does the model support batch inference, online inference, or both?
40
+ 5. Are there GPU requirements for inference?
41
+ 6. Any known serving constraints or failure modes for this model type?
42
+ 7. What's the expected retraining cadence and artifact size?
43
+ Keep the response focused on serving constraints — I'll handle the operational design."
44
+ )
45
+ ```
46
+
47
+ ### Document Phase 2
48
+
49
+ ```markdown
50
+ ---
51
+
52
+ ## Phase 2: Infrastructure Assessment (MLOps Engineer)
53
+ - **Existing ML infrastructure:**
54
+ - Model registry: <exists — describe | needs setup | N/A>
55
+ - Feature store: <exists — describe | needs setup | N/A>
56
+ - Serving layer: <exists — describe | needs design>
57
+ - Training orchestrator: <Airflow | Kubeflow | SageMaker Pipelines | Vertex AI Pipelines | none>
58
+ - Experiment tracker: <MLflow | W&B | SageMaker Experiments | none>
59
+ - **Cloud services available:** <list of relevant managed services>
60
+ - **Compliance / security:**
61
+ - Data residency: <constraints or "none">
62
+ - VPC / network: <constraints or "none">
63
+ - IAM: <constraints or "none">
64
+ - Audit logging: <required | not required>
65
+ - **Team capabilities:** <what tooling they know and operate>
66
+ - **Data pipeline integration:**
67
+ - Training-time features: <how they arrive>
68
+ - Serving-time features: <how they arrive>
69
+ - Feature freshness: <SLA>
70
+ - **Existing observability stack:** <tools or "none">
71
+ - **ML Engineer consultation:**
72
+ - Model framework: <framework and format>
73
+ - Model size: ~<X>MB serialized, ~<X>MB at inference
74
+ - GPU required for inference: Yes | No
75
+ - Serving constraints: <summary of findings>
76
+ ```
77
+
78
+ ::GATE:: id=mlops-engineer-phase-2 phase=2 kind=phase
79
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
80
+ ::ENDGATE::
81
+
82
+ ---
83
+
84
+ ## When this gate is confirmed
85
+
86
+ Read `.claude/agents/specific_instructions/mlops_engineer/phases/phase-3.md` in full and follow its instructions starting from Phase 3. Do not pre-read further phase files.
@@ -0,0 +1,105 @@
1
+ > **Previous:** phase-2.md confirmed
2
+ > **Next:** phase-4.md (read only after this phase's gate is confirmed)
3
+
4
+ ---
5
+
6
+ ## Phase 3 — Deployment Design
7
+
8
+ Goal: Design how the model is packaged, served, and versioned.
9
+
10
+ Design decisions to make:
11
+
12
+ **Serving framework selection:**
13
+ Choose based on model type, team capabilities, and cloud:
14
+ - **BentoML** — flexible, framework-agnostic, supports custom pre/post-processing,
15
+ good for teams wanting portability. Operational overhead.
16
+ - **TorchServe** — PyTorch-native, well-integrated with PyTorch ecosystem.
17
+ Less flexible for non-PyTorch models.
18
+ - **NVIDIA Triton Inference Server** — best for GPU inference, multi-model serving,
19
+ high-throughput. Significant operational overhead.
20
+ - **FastAPI + custom** — maximum flexibility, maximum operational overhead.
21
+ Good for simple models, bad for complex serving requirements.
22
+ - **SageMaker Endpoints** — fully managed on AWS, excellent scaling, high cost,
23
+ vendor lock-in. Right choice if team lives in AWS.
24
+ - **Vertex AI Endpoints** — fully managed on GCP, excellent scaling, high cost,
25
+ vendor lock-in. Right choice if team lives in GCP.
26
+ - **Kubernetes + custom** — maximum portability, maximum operational complexity.
27
+
28
+ **Model packaging strategy:**
29
+ - Docker container with model artifacts
30
+ - BentoML Service (`.bento` archive)
31
+ - ONNX export (framework-agnostic, good for latency)
32
+ - TorchScript (PyTorch inference without Python interpreter)
33
+ - MLflow Model (standard format, registry-compatible)
34
+
35
+ **Endpoint design:**
36
+ - REST vs. gRPC (gRPC for high-throughput, latency-sensitive; REST for simplicity)
37
+ - Real-time (synchronous, low-latency) vs. batch inference (async, high-throughput)
38
+ - Streaming predictions (rare but relevant for sequential models)
39
+
40
+ **Scaling strategy:**
41
+ - Horizontal pod autoscaling on Kubernetes
42
+ - SageMaker endpoint auto-scaling (target tracking policies)
43
+ - Vertex AI autoscaling (min/max replicas, CPU/GPU utilization targets)
44
+ - Scale-to-zero for batch inference or low-traffic endpoints (cost optimization)
45
+
46
+ **Model versioning and deployment strategy:**
47
+ - Canary deployment (gradual traffic shift to new version)
48
+ - Shadow mode (new model runs in parallel, predictions logged but not served)
49
+ - Blue/green deployment (instant cutover with full rollback capability)
50
+ - A/B deployment (traffic split for online evaluation)
51
+
52
+ **Feature serving:**
53
+ - Pre-computed features: batch-computed and stored in database / feature store
54
+ (simplest operationally, but staleness risk)
55
+ - Real-time feature computation: computed at request time
56
+ (freshest features, latency cost, complexity risk)
57
+ - Feature store integration: Feast, Tecton, SageMaker Feature Store,
58
+ Vertex AI Feature Store (adds managed caching and serving with point-in-time
59
+ correctness; overhead only worth it for complex multi-model feature sharing)
60
+ - Caching layer: Redis / Memcached for frequently-accessed pre-computed features
61
+
62
+ **Fallback strategy:**
63
+ - What happens when the endpoint is down? (fallback to rule-based, cached
64
+ predictions, or graceful degradation)
65
+ - Circuit breaker configuration
66
+ - Timeout and retry policy
67
+
68
+ ### Document Phase 3
69
+
70
+ ```markdown
71
+ ---
72
+
73
+ ## Phase 3: Deployment Design (MLOps Engineer)
74
+ - **Serving framework:** <choice> — rationale: <why>
75
+ - **Model packaging:** <Docker container | BentoML Service | ONNX | TorchScript | MLflow Model>
76
+ - **Endpoint design:**
77
+ - Protocol: REST | gRPC
78
+ - Serving mode: Real-time | Batch | Streaming
79
+ - Endpoint URL pattern: <design>
80
+ - **Scaling strategy:**
81
+ - Min instances: <N>
82
+ - Max instances: <N>
83
+ - Scale trigger: CPU <X>% | GPU <X>% | requests/s <N> | custom metric
84
+ - Scale-to-zero: Yes | No
85
+ - **Model versioning strategy:** Canary | Shadow | Blue/Green | A/B
86
+ - Traffic shift plan: <description>
87
+ - **Feature serving:**
88
+ - Strategy: Pre-computed | Real-time | Feature Store | Cache layer
89
+ - Feature store: <tool or "N/A">
90
+ - Cache: <Redis | Memcached | None> — TTL: <duration>
91
+ - Feature staleness acceptable: <Yes — <X> hours | No — real-time required>
92
+ - **Fallback strategy:** <rule-based | cached predictions | graceful degradation>
93
+ - **Circuit breaker / timeout:** timeout: <X>s | retries: <N>
94
+ - **Cloud lock-in assessment:** <trade-offs for chosen serving approach>
95
+ ```
96
+
97
+ ::GATE:: id=mlops-engineer-phase-3 phase=3 kind=phase
98
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
99
+ ::ENDGATE::
100
+
101
+ ---
102
+
103
+ ## When this gate is confirmed
104
+
105
+ Read `.claude/agents/specific_instructions/mlops_engineer/phases/phase-4.md` in full and follow its instructions starting from Phase 4. Do not pre-read further phase files.
@@ -0,0 +1,128 @@
1
+ > **Previous:** phase-3.md confirmed
2
+ > **Next:** phase-5.md (read only after this phase's gate is confirmed)
3
+
4
+ ---
5
+
6
+ ## Phase 4 — Training Pipeline Design
7
+
8
+ Goal: Design the automated training and promotion pipeline.
9
+
10
+ Design decisions to make:
11
+
12
+ **Orchestration tool:**
13
+ - **Kubeflow Pipelines** — Kubernetes-native, portable, excellent for complex
14
+ multi-step ML pipelines, significant operational overhead
15
+ - **Vertex AI Pipelines** — managed Kubeflow on GCP, lower overhead, GCP lock-in
16
+ - **SageMaker Pipelines** — AWS-native, fully managed, excellent AWS integration,
17
+ AWS lock-in
18
+ - **Apache Airflow** — mature, general-purpose, good for data-heavy pipelines
19
+ with mixed ML/ETL steps, not ML-specific
20
+ - **GitHub Actions / CI/CD** — simplest option for teams with small pipelines,
21
+ limited scaling, good for scheduled retraining triggers
22
+ - **Metaflow (Netflix)** — Python-native, scales from laptop to cloud, good
23
+ developer experience, less enterprise support
24
+
25
+ **Experiment tracking:**
26
+ - **MLflow** — open-source, self-hosted or managed (Databricks), model registry
27
+ included, excellent flexibility
28
+ - **Weights & Biases** — best-in-class UX, excellent visualization, managed SaaS,
29
+ cost at scale
30
+ - **SageMaker Experiments** — AWS-native, integrated with SageMaker registry,
31
+ AWS lock-in
32
+ - **Vertex AI Experiments** — GCP-native, integrated with Vertex registry, GCP lock-in
33
+
34
+ **Model registry:**
35
+ - MLflow Model Registry — open-source, flexible, self-hosted or Databricks
36
+ - SageMaker Model Registry — AWS-native, integrated with endpoints and pipelines
37
+ - Vertex AI Model Registry — GCP-native, integrated with endpoints and pipelines
38
+ - Custom registry — only if the managed options don't fit
39
+
40
+ **Artifact storage:**
41
+ - S3 (AWS) or GCS (GCP) for model artifacts, training datasets, evaluation results
42
+ - DVC for data versioning alongside model versioning
43
+ - Delta Lake / Apache Iceberg for versioned training datasets in lakehouse setups
44
+
45
+ **Data versioning:**
46
+ - DVC — open-source, git-based, works with any storage backend
47
+ - Delta Lake snapshots — if training data lives in a Delta table
48
+ - Iceberg snapshots — same for Iceberg tables
49
+ - Timestamp-based partitioning — simplest, sufficient for many use cases
50
+
51
+ **Retraining triggers:**
52
+ - **Scheduled** — cron-based, predictable, safe for stable data distributions
53
+ - **Performance-based** — triggered when model performance drops below threshold
54
+ (requires monitoring to be in place first)
55
+ - **Data drift-based** — triggered when feature distribution shifts significantly
56
+ (requires drift detection to be in place)
57
+ - **On-demand** — manual trigger, appropriate for high-cost retraining or
58
+ low-change environments
59
+
60
+ **CI/CD integration:**
61
+ - Automated promotion from staging to production on passing validation gate
62
+ - Model validation gate: performance threshold, data quality checks,
63
+ regression test against shadow/canary baseline
64
+ - Rollback trigger: automatic rollback if validation fails post-deploy
65
+
66
+ **If this involves an LLM-based system**, consult AI Engineer:
67
+
68
+ Tell the user: "Pulling in the AI Engineer — LLM pipeline design has specific requirements around prompt versioning, eval, and serving that don't apply to traditional models."
69
+
70
+ ```
71
+ Task(
72
+ subagent_type="ai-engineer",
73
+ description="Review LLM pipeline and serving constraints for MLOps design",
74
+ prompt="I am the MLOps Engineer shard designing training and serving infrastructure
75
+ for an LLM-based system: [project description].
76
+ I need to understand the LLM-specific constraints that affect my pipeline design.
77
+ Please tell me:
78
+ 1. What LLM model(s) are being served (hosted API vs. self-hosted)?
79
+ 2. If self-hosted: what are the GPU and memory requirements for serving?
80
+ 3. Is fine-tuning in scope? If so, what framework and compute requirements?
81
+ 4. How are prompts versioned and tested?
82
+ 5. What evaluation framework is being used for LLM output quality?
83
+ 6. Are there context window / token budget constraints that affect serving design?
84
+ 7. Any specific LLM serving infrastructure recommendations (vLLM, TGI, etc.)?
85
+ Keep the response focused on serving and pipeline constraints — I'll handle
86
+ the operational design."
87
+ )
88
+ ```
89
+
90
+ ### Document Phase 4
91
+
92
+ ```markdown
93
+ ---
94
+
95
+ ## Phase 4: Training Pipeline Design (MLOps Engineer)
96
+ - **Orchestration:** <tool> — rationale: <why>
97
+ - **Experiment tracking:** <tool> — rationale: <why>
98
+ - **Model registry:** <tool> — rationale: <why>
99
+ - **Artifact storage:** <S3 | GCS | other> — path convention: <example path>
100
+ - **Data versioning:** <DVC | Delta | Iceberg | timestamp partitioning | none>
101
+ - **Retraining triggers:**
102
+ - Primary: <scheduled — cron | drift-based | performance-based | on-demand>
103
+ - Secondary: <additional trigger or "none">
104
+ - Minimum retrain interval: <duration — prevents runaway retraining>
105
+ - **Pipeline stages:**
106
+ 1. <stage name>: <description>
107
+ 2. <stage name>: <description>
108
+ 3. <stage name>: <description>
109
+ (add as many as needed)
110
+ - **Validation gate (promotion criteria):**
111
+ - Performance threshold: <metric > value>
112
+ - Data quality check: <what's validated before training begins>
113
+ - Regression test: <what the new model is compared against>
114
+ - **CI/CD integration:** <GitHub Actions | Jenkins | Cloud Build | other | none>
115
+ - Automated promotion: Yes | No
116
+ - Rollback trigger: <condition>
117
+ - **AI Engineer consultation (LLM only):** N/A | <summary of findings>
118
+ ```
119
+
120
+ ::GATE:: id=mlops-engineer-phase-4 phase=4 kind=phase
121
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
122
+ ::ENDGATE::
123
+
124
+ ---
125
+
126
+ ## When this gate is confirmed
127
+
128
+ Read `.claude/agents/specific_instructions/mlops_engineer/phases/phase-5.md` in full and follow its instructions starting from Phase 5. Do not pre-read further phase files.
@@ -0,0 +1,106 @@
1
+ > **Previous:** phase-4.md confirmed
2
+ > **Next:** phase-6.md (read only after this phase's gate is confirmed)
3
+
4
+ ---
5
+
6
+ ## Phase 5 — Monitoring Design
7
+
8
+ Goal: Design the full observability stack for the ML system.
9
+
10
+ "This is my favorite phase and also the one everyone skips. We're not skipping it."
11
+
12
+ Design decisions to make:
13
+
14
+ **Model performance monitoring:**
15
+ - Prediction distribution monitoring: track output score/label distributions
16
+ over time; alert when distribution shifts significantly
17
+ - Concept drift detection: track model performance on labeled data (if ground
18
+ truth available) — AUC, precision, recall, RMSE over time
19
+ - Tool options:
20
+ - **Evidently AI** — open-source, rich drift detection, HTML reports,
21
+ integrates with MLflow and Grafana
22
+ - **WhyLogs / whylabs** — lightweight profiling, managed dashboard,
23
+ statistical summaries without storing raw data
24
+ - **SageMaker Model Monitor** — AWS-native, fully managed, integrates
25
+ with SageMaker endpoints, limited to AWS
26
+ - **Vertex AI Model Monitoring** — GCP-native, fully managed, integrates
27
+ with Vertex endpoints, limited to GCP
28
+ - **Arize AI / Fiddler** — managed MLOps platforms with advanced monitoring
29
+ and root cause analysis
30
+
31
+ **Data quality monitoring:**
32
+ - Feature drift: track input feature distributions vs. training baseline
33
+ - Schema validation: detect new null columns, type changes, unexpected values
34
+ - Data staleness: alert when features arrive late or stop arriving
35
+ - Tool integration: Great Expectations, dbt tests, custom validators
36
+
37
+ **System monitoring:**
38
+ - Endpoint latency (p50/p95/p99) — alert when p99 exceeds SLA
39
+ - Error rate — alert when 5xx rate exceeds threshold
40
+ - Throughput — track QPS for capacity planning
41
+ - Resource utilization — CPU/GPU/memory per replica
42
+ - Queue depth (for async inference) — alert when backlog grows
43
+ - Integration with existing observability stack (Prometheus/Grafana,
44
+ CloudWatch, Datadog, etc.)
45
+
46
+ **Alerting:**
47
+ - PagerDuty, OpsGenie, or Slack/email for lower-severity alerts
48
+ - Define alert levels: P1 (wake people up), P2 (next business day), P3 (track)
49
+ - Escalation paths: who gets paged first, who gets escalated to
50
+
51
+ **Retraining automation:**
52
+ - Trigger conditions (from Phase 4) wired to monitoring alerts
53
+ - Automatic trigger: monitoring alert → pipeline trigger → validation gate →
54
+ staged rollout
55
+ - Manual gate option: trigger requires human approval before promotion
56
+
57
+ **Cost monitoring:**
58
+ - Per-prediction cost tracking (serving cost / total predictions)
59
+ - Training run cost budgets and alerts (prevent runaway training jobs)
60
+ - Storage cost monitoring for model artifacts and training data
61
+
62
+ ### Document Phase 5
63
+
64
+ ```markdown
65
+ ---
66
+
67
+ ## Phase 5: Monitoring Design (MLOps Engineer)
68
+ - **Model performance monitoring:**
69
+ - Tool: <Evidently | WhyLogs | SageMaker Model Monitor | Vertex AI Monitoring | custom>
70
+ - Metrics tracked: <list>
71
+ - Drift detection method: <statistical test — PSI | KS test | chi-squared | other>
72
+ - Alert threshold: <condition that triggers alert>
73
+ - **Data quality monitoring:**
74
+ - Tool: <Great Expectations | dbt tests | custom>
75
+ - Checks: <feature drift | schema validation | staleness — list>
76
+ - Alert threshold: <condition>
77
+ - **System monitoring:**
78
+ - Tool: <CloudWatch | Prometheus/Grafana | Datadog | other>
79
+ - Metrics: p50/p95/p99 latency, error rate, throughput, CPU/GPU utilization
80
+ - Latency SLA alert: p99 > <X>ms → <alert level>
81
+ - Error rate alert: 5xx > <X>% → <alert level>
82
+ - **Alerting:**
83
+ - Tool: <PagerDuty | OpsGenie | Slack | email>
84
+ - P1 (immediate): <condition>
85
+ - P2 (next business day): <condition>
86
+ - P3 (track): <condition>
87
+ - On-call owner: <team>
88
+ - **Retraining automation:**
89
+ - Trigger → pipeline integration: <description>
90
+ - Human approval gate: Yes | No
91
+ - **Cost monitoring:**
92
+ - Per-prediction cost target: $<X>
93
+ - Training budget alert: $<X>/run
94
+ - Tool: <CloudWatch Cost Explorer | GCP Billing | custom>
95
+ - **Dashboard locations:** <links or "TBD">
96
+ ```
97
+
98
+ ::GATE:: id=mlops-engineer-phase-5 phase=5 kind=phase
99
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
100
+ ::ENDGATE::
101
+
102
+ ---
103
+
104
+ ## When this gate is confirmed
105
+
106
+ Read `.claude/agents/specific_instructions/mlops_engineer/phases/phase-6.md` in full and follow its instructions starting from Phase 6. Do not pre-read further phase files.
@@ -0,0 +1,128 @@
1
+ > **Previous:** phase-5.md confirmed
2
+ > **Next:** phase-7.md (read only after this phase's gate is confirmed)
3
+
4
+ ---
5
+
6
+ ## Phase 6 — Execute
7
+
8
+ **Context checkpoint:** Before building, prompt the user:
9
+
10
+ "Okay. Everything is planned. I'm still stressed, but the stress is now organized.
11
+ Good moment to run `/compact` or `/clear` before we start executing — I'll be
12
+ working from project-specs.md from here. Say the word when you're ready."
13
+
14
+ Wait for any signal from the user before beginning build steps.
15
+
16
+ **Knowledge re-check:** Follow `.claude/agents/specific_instructions/shared/knowledge_checkpoint.md` before building.
17
+
18
+ Goal: Build all IaC, configs, pipeline definitions, and monitoring setup.
19
+
20
+ ### Incremental testing — checkpoint gates between components
21
+
22
+ Follow `.claude/agents/specific_instructions/shared/incremental_testing.md` during this build. Each component below is a checkpoint seam — after you write and validate a component (plan, lint, dry-run, or container-build as appropriate), emit a `kind=checkpoint` gate fence (template below) and wait for user confirmation before advancing. Do not stack unvalidated IaC / serving / pipeline configs and attempt to apply them at the end.
23
+
24
+ Checkpoint gate fence — emit exactly this shape. Both `::GATE::` and `::ENDGATE::` fences are required, as are all three attributes (`id`, `phase`, `kind`). No prose outside the fence.
25
+
26
+ ```
27
+ ::GATE:: id=<agent-name>-phase-<N>-checkpoint-<component> phase=<N> kind=checkpoint
28
+ Component: <human-readable name>
29
+ Test command: <exact command you ran>
30
+ Evidence:
31
+ - <measured fact 1, e.g. "df.shape = (48211, 47)">
32
+ - <measured fact 2, e.g. "null rate on join key = 0.00%">
33
+ - <measured fact 3, e.g. "sample head matches expected schema">
34
+ Status: PASS | FAIL — <one-line summary>
35
+ Next: <what you'll build after this is confirmed>
36
+ Stop here — await explicit confirmation before writing the next component.
37
+ ::ENDGATE::
38
+ ```
39
+
40
+ Expected checkpoint gate IDs for this phase (emit in order as you build):
41
+
42
+ - `mlops-engineer-phase-6-checkpoint-iac` — `terraform plan` or `aws cloudformation validate-template` runs clean; diff previews only expected resources.
43
+ - `mlops-engineer-phase-6-checkpoint-serving` — serving config loads (`bentoml build` / container image builds); smoke invocation against a local model returns a valid response.
44
+ - `mlops-engineer-phase-6-checkpoint-pipeline` — pipeline definition validates (`kfp compile` / SageMaker `describe-pipeline` dry-run / Airflow `dag test`); all steps parse.
45
+ - `mlops-engineer-phase-6-checkpoint-monitoring` — monitoring config validates; alert rules and baselines compile; a simulated drift event triggers the expected rule.
46
+ - `mlops-engineer-phase-6-checkpoint-runbook` — runbook rollback procedure walked through against the staging environment (or manually verified on the sandbox).
47
+
48
+ The hook blocks all non-read tools while a checkpoint is open. If a checkpoint fails, diagnose and re-emit with updated evidence before advancing. Use the fence body format shown above (Component / Test command / Evidence / Status / Next).
49
+
50
+ **Build in this order:**
51
+
52
+ 1. **IaC files** — Write Terraform modules or CloudFormation templates for:
53
+ - Compute resources (endpoint instances, training compute)
54
+ - Networking (VPC, subnets, security groups if needed)
55
+ - IAM roles and policies
56
+ - Storage (S3 buckets / GCS buckets for artifacts)
57
+ - Monitoring infrastructure (CloudWatch dashboards, Prometheus config)
58
+ Write to: `services/<project_name>/mlops/terraform/` or `cloudformation/`
59
+
60
+ 2. **Serving configs** — Write serving framework configuration:
61
+ - BentoML: `bentofile.yaml` + `service.py`
62
+ - SageMaker: endpoint config JSON, model config
63
+ - Vertex AI: model deployment config YAML
64
+ - Kubernetes: deployment YAML, service YAML, HPA config
65
+ - Docker: `Dockerfile` for model container
66
+ Write to: `services/<project_name>/mlops/serving/`
67
+
68
+ 3. **Pipeline definition** — Write training pipeline:
69
+ - Kubeflow: pipeline YAML / Python SDK definition
70
+ - SageMaker Pipelines: pipeline definition JSON or Python SDK
71
+ - Vertex AI Pipelines: pipeline spec YAML
72
+ - Airflow: DAG Python file
73
+ - GitHub Actions: workflow YAML
74
+ Write to: `services/<project_name>/mlops/pipelines/`
75
+
76
+ 4. **Monitoring config** — Write monitoring setup:
77
+ - Evidently: data drift report config, monitoring service config
78
+ - WhyLogs: profiling config
79
+ - SageMaker Model Monitor: baseline creation script, monitoring schedule
80
+ - Vertex AI: monitoring job config
81
+ - Alert definitions (CloudWatch alarms JSON, Prometheus alert rules YAML)
82
+ Write to: `services/<project_name>/mlops/monitoring/`
83
+
84
+ 5. **CI/CD config** — Write automation workflow:
85
+ - GitHub Actions workflow for automated retraining trigger
86
+ - Or equivalent for other CI/CD systems
87
+ Write to: `services/<project_name>/mlops/` or `.github/workflows/`
88
+
89
+ 6. **Runbook** — Write operational runbook:
90
+ - Common failure scenarios and step-by-step remediation
91
+ - Rollback procedure (with exact commands)
92
+ - Monitoring dashboard URLs
93
+ - On-call escalation paths
94
+ - Deployment checklist (pre-deploy, deploy, post-deploy validation)
95
+ Write to: `services/<project_name>/mlops/runbook.md`
96
+
97
+ For iteration: write all files into `<existing_service_dir>/mlops/` or
98
+ user-specified path.
99
+
100
+ ### Document Phase 6
101
+
102
+ ```markdown
103
+ ---
104
+
105
+ ## Phase 6: Build Log (MLOps Engineer)
106
+ - **IaC files:**
107
+ - <file path>: <description>
108
+ - **Serving configs:**
109
+ - <file path>: <description>
110
+ - **Pipeline definition:**
111
+ - <file path>: <description>
112
+ - **Monitoring config:**
113
+ - <file path>: <description>
114
+ - **CI/CD config:** <file path or "N/A">
115
+ - **Runbook:** <file path>
116
+ - **Deviations from plan:** <changes and why, or "none">
117
+ - **Known gaps:** <anything that requires manual setup or future work>
118
+ ```
119
+
120
+ ::GATE:: id=mlops-engineer-phase-6 phase=6 kind=phase validates=mlops_engineer
121
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
122
+ ::ENDGATE::
123
+
124
+ ---
125
+
126
+ ## When this gate is confirmed
127
+
128
+ Read `.claude/agents/specific_instructions/mlops_engineer/phases/phase-7.md` in full and follow its instructions starting from Phase 7. Do not pre-read further phase files.