@proflandrigan/shards 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (397) hide show
  1. package/README.md +475 -0
  2. package/package.json +37 -0
  3. package/src/agents/academic.md +276 -0
  4. package/src/agents/ai-engineer.md +377 -0
  5. package/src/agents/analytics-engineer.md +364 -0
  6. package/src/agents/applied-ml-scientist.md +410 -0
  7. package/src/agents/backend-engineer.md +255 -0
  8. package/src/agents/bi-engineer.md +333 -0
  9. package/src/agents/data-analyst.md +343 -0
  10. package/src/agents/data-engineer.md +260 -0
  11. package/src/agents/data-modeller.md +386 -0
  12. package/src/agents/data-scientist.md +366 -0
  13. package/src/agents/deep-learning-engineer.md +389 -0
  14. package/src/agents/ml-engineer.md +424 -0
  15. package/src/agents/mlops-engineer.md +339 -0
  16. package/src/agents/researcher.md +187 -0
  17. package/src/agents/specific_instructions/academic/critical_review.md +263 -0
  18. package/src/agents/specific_instructions/academic/report.md +113 -0
  19. package/src/agents/specific_instructions/ai_engineer/advise.md +162 -0
  20. package/src/agents/specific_instructions/ai_engineer/bi_engineer_handoff.md +86 -0
  21. package/src/agents/specific_instructions/ai_engineer/experiment.md +471 -0
  22. package/src/agents/specific_instructions/ai_engineer/experiment_ui_mode.md +44 -0
  23. package/src/agents/specific_instructions/ai_engineer/phases/index.md +45 -0
  24. package/src/agents/specific_instructions/ai_engineer/phases/phase-1.md +55 -0
  25. package/src/agents/specific_instructions/ai_engineer/phases/phase-2.md +86 -0
  26. package/src/agents/specific_instructions/ai_engineer/phases/phase-3.md +96 -0
  27. package/src/agents/specific_instructions/ai_engineer/phases/phase-4.md +138 -0
  28. package/src/agents/specific_instructions/ai_engineer/phases/phase-5.md +157 -0
  29. package/src/agents/specific_instructions/ai_engineer/phases/phase-6.md +196 -0
  30. package/src/agents/specific_instructions/ai_engineer/phases/phase-7.md +313 -0
  31. package/src/agents/specific_instructions/ai_engineer/phases.md +1011 -0
  32. package/src/agents/specific_instructions/ai_engineer/prompt_lab.md +161 -0
  33. package/src/agents/specific_instructions/ai_engineer/prompt_lab_ui_mode.md +28 -0
  34. package/src/agents/specific_instructions/ai_engineer/research.md +393 -0
  35. package/src/agents/specific_instructions/ai_engineer/research_ui_mode.md +66 -0
  36. package/src/agents/specific_instructions/ai_engineer/review.md +159 -0
  37. package/src/agents/specific_instructions/ai_engineer/validation_checklist.md +182 -0
  38. package/src/agents/specific_instructions/analytics_engineer/advise.md +155 -0
  39. package/src/agents/specific_instructions/analytics_engineer/bi_engineer_handoff.md +91 -0
  40. package/src/agents/specific_instructions/analytics_engineer/data_analyst_handoff.md +84 -0
  41. package/src/agents/specific_instructions/analytics_engineer/deep_phases.md +818 -0
  42. package/src/agents/specific_instructions/analytics_engineer/phases_deep/index.md +24 -0
  43. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-1.md +77 -0
  44. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-2.md +106 -0
  45. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-3.md +93 -0
  46. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-4.md +79 -0
  47. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-5.md +61 -0
  48. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-6.md +45 -0
  49. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-7.md +235 -0
  50. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-8.md +221 -0
  51. package/src/agents/specific_instructions/analytics_engineer/phases_quick/index.md +19 -0
  52. package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-1.md +47 -0
  53. package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-2.md +78 -0
  54. package/src/agents/specific_instructions/analytics_engineer/quick_phases.md +112 -0
  55. package/src/agents/specific_instructions/analytics_engineer/review.md +167 -0
  56. package/src/agents/specific_instructions/analytics_engineer/service_mode.md +369 -0
  57. package/src/agents/specific_instructions/analytics_engineer/ui_mode.md +45 -0
  58. package/src/agents/specific_instructions/analytics_engineer/update.md +162 -0
  59. package/src/agents/specific_instructions/analytics_engineer/validation_checklist.md +121 -0
  60. package/src/agents/specific_instructions/applied_ml_scientist/advise.md +143 -0
  61. package/src/agents/specific_instructions/applied_ml_scientist/phases/index.md +21 -0
  62. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-1.md +51 -0
  63. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-2.md +66 -0
  64. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-3.md +113 -0
  65. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-4.md +104 -0
  66. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-5.md +156 -0
  67. package/src/agents/specific_instructions/applied_ml_scientist/phases.md +428 -0
  68. package/src/agents/specific_instructions/applied_ml_scientist/research.md +379 -0
  69. package/src/agents/specific_instructions/applied_ml_scientist/review.md +142 -0
  70. package/src/agents/specific_instructions/applied_ml_scientist/validation_checklist.md +136 -0
  71. package/src/agents/specific_instructions/backend_engineer/clean.md +149 -0
  72. package/src/agents/specific_instructions/backend_engineer/review.md +91 -0
  73. package/src/agents/specific_instructions/backend_engineer/review_checklist.md +54 -0
  74. package/src/agents/specific_instructions/backend_engineer/service_mode.md +67 -0
  75. package/src/agents/specific_instructions/bi_engineer/advise.md +137 -0
  76. package/src/agents/specific_instructions/bi_engineer/data_analyst_handoff.md +77 -0
  77. package/src/agents/specific_instructions/bi_engineer/incoming_handoff.md +45 -0
  78. package/src/agents/specific_instructions/bi_engineer/phases/index.md +20 -0
  79. package/src/agents/specific_instructions/bi_engineer/phases/phase-1.md +164 -0
  80. package/src/agents/specific_instructions/bi_engineer/phases/phase-2.md +92 -0
  81. package/src/agents/specific_instructions/bi_engineer/phases/phase-3.md +121 -0
  82. package/src/agents/specific_instructions/bi_engineer/phases/phase-4.md +106 -0
  83. package/src/agents/specific_instructions/bi_engineer/phases.md +451 -0
  84. package/src/agents/specific_instructions/bi_engineer/review.md +166 -0
  85. package/src/agents/specific_instructions/bi_engineer/update.md +147 -0
  86. package/src/agents/specific_instructions/bi_engineer/validation_checklist.md +124 -0
  87. package/src/agents/specific_instructions/data_analyst/advise.md +138 -0
  88. package/src/agents/specific_instructions/data_analyst/explain.md +221 -0
  89. package/src/agents/specific_instructions/data_analyst/incoming_handoff.md +40 -0
  90. package/src/agents/specific_instructions/data_analyst/phases/index.md +20 -0
  91. package/src/agents/specific_instructions/data_analyst/phases/phase-1.md +159 -0
  92. package/src/agents/specific_instructions/data_analyst/phases/phase-2.md +112 -0
  93. package/src/agents/specific_instructions/data_analyst/phases/phase-3.md +265 -0
  94. package/src/agents/specific_instructions/data_analyst/phases/phase-4.md +100 -0
  95. package/src/agents/specific_instructions/data_analyst/phases.md +501 -0
  96. package/src/agents/specific_instructions/data_analyst/review.md +138 -0
  97. package/src/agents/specific_instructions/data_analyst/ui_mode.md +26 -0
  98. package/src/agents/specific_instructions/data_analyst/update.md +144 -0
  99. package/src/agents/specific_instructions/data_analyst/validation_checklist.md +95 -0
  100. package/src/agents/specific_instructions/data_engineer/advise.md +137 -0
  101. package/src/agents/specific_instructions/data_engineer/phases.md +466 -0
  102. package/src/agents/specific_instructions/data_engineer/phases_deep/index.md +23 -0
  103. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-1.md +49 -0
  104. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-2.md +93 -0
  105. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-3.md +55 -0
  106. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-4.md +48 -0
  107. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-5.md +40 -0
  108. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-6.md +102 -0
  109. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-7.md +87 -0
  110. package/src/agents/specific_instructions/data_engineer/phases_quick/index.md +19 -0
  111. package/src/agents/specific_instructions/data_engineer/phases_quick/phase-1.md +45 -0
  112. package/src/agents/specific_instructions/data_engineer/phases_quick/phase-2.md +54 -0
  113. package/src/agents/specific_instructions/data_engineer/review.md +135 -0
  114. package/src/agents/specific_instructions/data_engineer/validation_checklist.md +136 -0
  115. package/src/agents/specific_instructions/data_modeller/advise.md +137 -0
  116. package/src/agents/specific_instructions/data_modeller/phases.md +581 -0
  117. package/src/agents/specific_instructions/data_modeller/phases_deep/index.md +23 -0
  118. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-1.md +52 -0
  119. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-2.md +113 -0
  120. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-3.md +47 -0
  121. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-4.md +51 -0
  122. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-5.md +45 -0
  123. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-6.md +105 -0
  124. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-7.md +136 -0
  125. package/src/agents/specific_instructions/data_modeller/phases_quick/index.md +19 -0
  126. package/src/agents/specific_instructions/data_modeller/phases_quick/phase-1.md +47 -0
  127. package/src/agents/specific_instructions/data_modeller/phases_quick/phase-2.md +65 -0
  128. package/src/agents/specific_instructions/data_modeller/review.md +141 -0
  129. package/src/agents/specific_instructions/data_modeller/service_mode.md +218 -0
  130. package/src/agents/specific_instructions/data_modeller/validation_checklist.md +125 -0
  131. package/src/agents/specific_instructions/data_scientist/advise.md +158 -0
  132. package/src/agents/specific_instructions/data_scientist/bi_engineer_handoff.md +63 -0
  133. package/src/agents/specific_instructions/data_scientist/experiment.md +482 -0
  134. package/src/agents/specific_instructions/data_scientist/experiment_ui_mode.md +44 -0
  135. package/src/agents/specific_instructions/data_scientist/explain.md +247 -0
  136. package/src/agents/specific_instructions/data_scientist/greenfield_data.md +35 -0
  137. package/src/agents/specific_instructions/data_scientist/ml_engineer_handoff.md +52 -0
  138. package/src/agents/specific_instructions/data_scientist/notebook_walkthrough.md +76 -0
  139. package/src/agents/specific_instructions/data_scientist/phases/index.md +24 -0
  140. package/src/agents/specific_instructions/data_scientist/phases/phase-1.md +45 -0
  141. package/src/agents/specific_instructions/data_scientist/phases/phase-2.md +67 -0
  142. package/src/agents/specific_instructions/data_scientist/phases/phase-3.md +89 -0
  143. package/src/agents/specific_instructions/data_scientist/phases/phase-4.md +143 -0
  144. package/src/agents/specific_instructions/data_scientist/phases/phase-5.md +71 -0
  145. package/src/agents/specific_instructions/data_scientist/phases/phase-6.md +239 -0
  146. package/src/agents/specific_instructions/data_scientist/phases/phase-7.md +207 -0
  147. package/src/agents/specific_instructions/data_scientist/phases.md +651 -0
  148. package/src/agents/specific_instructions/data_scientist/research.md +345 -0
  149. package/src/agents/specific_instructions/data_scientist/research_ui_mode.md +52 -0
  150. package/src/agents/specific_instructions/data_scientist/review.md +136 -0
  151. package/src/agents/specific_instructions/data_scientist/service_mode.md +247 -0
  152. package/src/agents/specific_instructions/data_scientist/validation_checklist.md +183 -0
  153. package/src/agents/specific_instructions/deep_learning_engineer/advise.md +145 -0
  154. package/src/agents/specific_instructions/deep_learning_engineer/phases/index.md +21 -0
  155. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-1.md +74 -0
  156. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-2.md +98 -0
  157. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-3.md +76 -0
  158. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-4.md +128 -0
  159. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-5.md +292 -0
  160. package/src/agents/specific_instructions/deep_learning_engineer/phases.md +567 -0
  161. package/src/agents/specific_instructions/deep_learning_engineer/research.md +389 -0
  162. package/src/agents/specific_instructions/deep_learning_engineer/review.md +155 -0
  163. package/src/agents/specific_instructions/deep_learning_engineer/validation_checklist.md +147 -0
  164. package/src/agents/specific_instructions/ml_engineer/advise.md +174 -0
  165. package/src/agents/specific_instructions/ml_engineer/bi_engineer_handoff.md +71 -0
  166. package/src/agents/specific_instructions/ml_engineer/experiment.md +474 -0
  167. package/src/agents/specific_instructions/ml_engineer/experiment_ui_mode.md +44 -0
  168. package/src/agents/specific_instructions/ml_engineer/notebook_walkthrough.md +75 -0
  169. package/src/agents/specific_instructions/ml_engineer/phases/index.md +25 -0
  170. package/src/agents/specific_instructions/ml_engineer/phases/phase-1.md +49 -0
  171. package/src/agents/specific_instructions/ml_engineer/phases/phase-2.md +75 -0
  172. package/src/agents/specific_instructions/ml_engineer/phases/phase-3.md +124 -0
  173. package/src/agents/specific_instructions/ml_engineer/phases/phase-4.md +279 -0
  174. package/src/agents/specific_instructions/ml_engineer/phases/phase-5.md +160 -0
  175. package/src/agents/specific_instructions/ml_engineer/phases/phase-6-5.md +170 -0
  176. package/src/agents/specific_instructions/ml_engineer/phases/phase-6.md +295 -0
  177. package/src/agents/specific_instructions/ml_engineer/phases/phase-7.md +337 -0
  178. package/src/agents/specific_instructions/ml_engineer/phases.md +1068 -0
  179. package/src/agents/specific_instructions/ml_engineer/research.md +437 -0
  180. package/src/agents/specific_instructions/ml_engineer/research_ui_mode.md +71 -0
  181. package/src/agents/specific_instructions/ml_engineer/review.md +187 -0
  182. package/src/agents/specific_instructions/ml_engineer/service_mode.md +273 -0
  183. package/src/agents/specific_instructions/ml_engineer/validation_checklist.md +185 -0
  184. package/src/agents/specific_instructions/mlops_engineer/advise.md +139 -0
  185. package/src/agents/specific_instructions/mlops_engineer/phases/index.md +23 -0
  186. package/src/agents/specific_instructions/mlops_engineer/phases/phase-1.md +52 -0
  187. package/src/agents/specific_instructions/mlops_engineer/phases/phase-2.md +86 -0
  188. package/src/agents/specific_instructions/mlops_engineer/phases/phase-3.md +105 -0
  189. package/src/agents/specific_instructions/mlops_engineer/phases/phase-4.md +128 -0
  190. package/src/agents/specific_instructions/mlops_engineer/phases/phase-5.md +106 -0
  191. package/src/agents/specific_instructions/mlops_engineer/phases/phase-6.md +128 -0
  192. package/src/agents/specific_instructions/mlops_engineer/phases/phase-7.md +144 -0
  193. package/src/agents/specific_instructions/mlops_engineer/phases.md +671 -0
  194. package/src/agents/specific_instructions/mlops_engineer/review.md +164 -0
  195. package/src/agents/specific_instructions/mlops_engineer/service_mode.md +81 -0
  196. package/src/agents/specific_instructions/mlops_engineer/validation_checklist.md +151 -0
  197. package/src/agents/specific_instructions/researcher/critical_review.md +292 -0
  198. package/src/agents/specific_instructions/researcher/review_checklist.md +67 -0
  199. package/src/agents/specific_instructions/researcher/service_mode.md +224 -0
  200. package/src/agents/specific_instructions/shared/auto_verify_mode.md +141 -0
  201. package/src/agents/specific_instructions/shared/autonomous_research.md +1289 -0
  202. package/src/agents/specific_instructions/shared/behavioral_rules.md +36 -0
  203. package/src/agents/specific_instructions/shared/diverge_protocol.md +387 -0
  204. package/src/agents/specific_instructions/shared/engineering_guidelines.md +136 -0
  205. package/src/agents/specific_instructions/shared/experiment_versioning.md +184 -0
  206. package/src/agents/specific_instructions/shared/goal_mode.md +187 -0
  207. package/src/agents/specific_instructions/shared/incremental_testing.md +139 -0
  208. package/src/agents/specific_instructions/shared/intent_discovery.md +223 -0
  209. package/src/agents/specific_instructions/shared/join_path_protocol.md +168 -0
  210. package/src/agents/specific_instructions/shared/knowledge_checkpoint.md +83 -0
  211. package/src/agents/specific_instructions/shared/knowledge_harvest.md +220 -0
  212. package/src/agents/specific_instructions/shared/knowledge_retrieval.md +100 -0
  213. package/src/agents/specific_instructions/shared/notebook_walkthrough_protocol.md +367 -0
  214. package/src/agents/specific_instructions/shared/reviewer_verdict_protocol.md +74 -0
  215. package/src/agents/specific_instructions/shared/swarm_protocol.md +97 -0
  216. package/src/agents/specific_instructions/shared/validation_protocol.md +139 -0
  217. package/src/agents/specific_instructions/syn/arbiter.md +140 -0
  218. package/src/agents/specific_instructions/syn/brainstorm.md +550 -0
  219. package/src/agents/specific_instructions/syn/code_review.md +232 -0
  220. package/src/agents/specific_instructions/syn/diff.md +239 -0
  221. package/src/agents/specific_instructions/syn/final_review.md +65 -0
  222. package/src/agents/specific_instructions/syn/fixer.md +240 -0
  223. package/src/agents/specific_instructions/syn/free_form.md +130 -0
  224. package/src/agents/specific_instructions/syn/knowledge.md +468 -0
  225. package/src/agents/specific_instructions/syn/notebook_walkthrough.md +78 -0
  226. package/src/agents/specific_instructions/syn/panel_review.md +634 -0
  227. package/src/agents/specific_instructions/syn/pm.md +453 -0
  228. package/src/agents/specific_instructions/syn/pr_review.md +255 -0
  229. package/src/agents/specific_instructions/syn/slides.md +417 -0
  230. package/src/agents/syn.md +729 -0
  231. package/src/commands/academic.md +41 -0
  232. package/src/commands/ai-engineer.md +45 -0
  233. package/src/commands/analytics-engineer.md +48 -0
  234. package/src/commands/applied-ml-scientist.md +45 -0
  235. package/src/commands/backend-engineer.md +35 -0
  236. package/src/commands/bi-engineer.md +40 -0
  237. package/src/commands/brainstorm.md +24 -0
  238. package/src/commands/data-analyst.md +38 -0
  239. package/src/commands/data-engineer.md +37 -0
  240. package/src/commands/data-modeller.md +38 -0
  241. package/src/commands/data-scientist.md +38 -0
  242. package/src/commands/deep-learning-engineer.md +47 -0
  243. package/src/commands/end.md +49 -0
  244. package/src/commands/knowledge.md +24 -0
  245. package/src/commands/ml-engineer.md +42 -0
  246. package/src/commands/mlops-engineer.md +47 -0
  247. package/src/commands/notebook-walkthrough.md +58 -0
  248. package/src/commands/researcher.md +40 -0
  249. package/src/commands/resume.md +57 -0
  250. package/src/commands/review-pr.md +26 -0
  251. package/src/commands/shards-guide.md +41 -0
  252. package/src/commands/shards-ui.md +32 -0
  253. package/src/commands/shards.md +41 -0
  254. package/src/docs/01-getting-started/concepts.md +109 -0
  255. package/src/docs/01-getting-started/first-session.md +79 -0
  256. package/src/docs/01-getting-started/install.md +61 -0
  257. package/src/docs/02-agents/academic.md +71 -0
  258. package/src/docs/02-agents/ai-engineer.md +78 -0
  259. package/src/docs/02-agents/analytics-engineer.md +58 -0
  260. package/src/docs/02-agents/applied-ml-scientist.md +59 -0
  261. package/src/docs/02-agents/backend-engineer.md +58 -0
  262. package/src/docs/02-agents/bi-engineer.md +65 -0
  263. package/src/docs/02-agents/data-analyst.md +67 -0
  264. package/src/docs/02-agents/data-engineer.md +57 -0
  265. package/src/docs/02-agents/data-modeller.md +51 -0
  266. package/src/docs/02-agents/data-scientist.md +78 -0
  267. package/src/docs/02-agents/deep-learning-engineer.md +64 -0
  268. package/src/docs/02-agents/ml-engineer.md +80 -0
  269. package/src/docs/02-agents/mlops-engineer.md +59 -0
  270. package/src/docs/02-agents/overview.md +62 -0
  271. package/src/docs/02-agents/researcher.md +73 -0
  272. package/src/docs/02-agents/syn.md +88 -0
  273. package/src/docs/03-protocols/auto-verify.md +82 -0
  274. package/src/docs/03-protocols/autonomous-research.md +59 -0
  275. package/src/docs/03-protocols/behavioral-rules.md +35 -0
  276. package/src/docs/03-protocols/diverge.md +50 -0
  277. package/src/docs/03-protocols/engineering-guidelines.md +56 -0
  278. package/src/docs/03-protocols/experiment-versioning.md +38 -0
  279. package/src/docs/03-protocols/gate-pattern.md +65 -0
  280. package/src/docs/03-protocols/incremental-testing.md +68 -0
  281. package/src/docs/03-protocols/join-path.md +46 -0
  282. package/src/docs/03-protocols/knowledge-ledger.md +70 -0
  283. package/src/docs/03-protocols/reviewer-verdicts.md +39 -0
  284. package/src/docs/03-protocols/swarm.md +40 -0
  285. package/src/docs/03-protocols/validation.md +174 -0
  286. package/src/docs/04-ui/activity-bar.md +70 -0
  287. package/src/docs/04-ui/chat-pane.md +80 -0
  288. package/src/docs/04-ui/code-intel.md +62 -0
  289. package/src/docs/04-ui/file-editing.md +61 -0
  290. package/src/docs/04-ui/git.md +54 -0
  291. package/src/docs/04-ui/keybindings.md +79 -0
  292. package/src/docs/04-ui/knowledge-map.md +76 -0
  293. package/src/docs/04-ui/overview.md +93 -0
  294. package/src/docs/04-ui/panels.md +49 -0
  295. package/src/docs/04-ui/pinboard-selection.md +66 -0
  296. package/src/docs/04-ui/quick-open-palette.md +56 -0
  297. package/src/docs/04-ui/sessions.md +81 -0
  298. package/src/docs/04-ui/settings-permissions.md +56 -0
  299. package/src/docs/05-commands/reference.md +59 -0
  300. package/src/docs/06-outputs/directory-map.md +116 -0
  301. package/src/docs/07-workflows/ai-eval-first.md +57 -0
  302. package/src/docs/07-workflows/deep-study-to-production.md +76 -0
  303. package/src/docs/07-workflows/diverge-exploration.md +77 -0
  304. package/src/docs/07-workflows/quick-analysis.md +45 -0
  305. package/src/docs/08-integrations/claude-code-auto-mode.md +191 -0
  306. package/src/docs/08-integrations/google-slides.md +175 -0
  307. package/src/docs/README.md +30 -0
  308. package/src/docs/manifest.json +108 -0
  309. package/src/templates/analysis-template.md +20 -0
  310. package/src/templates/branch-report.md +46 -0
  311. package/src/templates/diff-report.md +88 -0
  312. package/src/templates/knowledge-index.md +7 -0
  313. package/src/templates/model-card-schema.json +186 -0
  314. package/src/templates/model-card-schema.md +88 -0
  315. package/src/templates/model-card.md +124 -0
  316. package/src/templates/project-plan.md +47 -0
  317. package/src/templates/project-specs.md +81 -0
  318. package/src/templates/report-template.md +43 -0
  319. package/src/templates/study-template.md +25 -0
  320. package/src/ui/cc-readonly.js +181 -0
  321. package/src/ui/chat-session.js +466 -0
  322. package/src/ui/css/base.css +136 -0
  323. package/src/ui/css/brainstorm.css +525 -0
  324. package/src/ui/css/chat.css +1405 -0
  325. package/src/ui/css/editor.css +546 -0
  326. package/src/ui/css/eval-dashboard.css +157 -0
  327. package/src/ui/css/experiment.css +237 -0
  328. package/src/ui/css/guide.css +186 -0
  329. package/src/ui/css/knowledge-map.css +383 -0
  330. package/src/ui/css/layout.css +431 -0
  331. package/src/ui/css/model-card.css +161 -0
  332. package/src/ui/css/notebook-walkthrough.css +271 -0
  333. package/src/ui/css/pr-review.css +403 -0
  334. package/src/ui/css/prompt-lab.css +325 -0
  335. package/src/ui/css/sessions.css +258 -0
  336. package/src/ui/css/sidebar.css +661 -0
  337. package/src/ui/css/terminal.css +113 -0
  338. package/src/ui/css/theme-light.css +542 -0
  339. package/src/ui/index.html +389 -0
  340. package/src/ui/js/agents.js +32 -0
  341. package/src/ui/js/bookmarks.js +230 -0
  342. package/src/ui/js/chat.js +1776 -0
  343. package/src/ui/js/code-intel.js +328 -0
  344. package/src/ui/js/command-palette.js +142 -0
  345. package/src/ui/js/events.js +591 -0
  346. package/src/ui/js/explorer.js +317 -0
  347. package/src/ui/js/file-view.js +477 -0
  348. package/src/ui/js/git.js +536 -0
  349. package/src/ui/js/guide.js +198 -0
  350. package/src/ui/js/hud.js +75 -0
  351. package/src/ui/js/init.js +351 -0
  352. package/src/ui/js/knowledge-map.js +906 -0
  353. package/src/ui/js/markdown.js +114 -0
  354. package/src/ui/js/monaco.js +164 -0
  355. package/src/ui/js/notebook-walkthrough.js +272 -0
  356. package/src/ui/js/notebook.js +448 -0
  357. package/src/ui/js/panels.js +2681 -0
  358. package/src/ui/js/pinboard.js +186 -0
  359. package/src/ui/js/quick-open.js +164 -0
  360. package/src/ui/js/selection-context.js +131 -0
  361. package/src/ui/js/sessions.js +256 -0
  362. package/src/ui/js/settings.js +476 -0
  363. package/src/ui/js/split-view.js +82 -0
  364. package/src/ui/js/state.js +343 -0
  365. package/src/ui/js/table.js +161 -0
  366. package/src/ui/js/tabs.js +284 -0
  367. package/src/ui/js/tabular.js +125 -0
  368. package/src/ui/js/terminal.js +354 -0
  369. package/src/ui/js/timeline.js +137 -0
  370. package/src/ui/js/utils.js +293 -0
  371. package/src/ui/notebook-kernel.py +790 -0
  372. package/src/ui/open-browser.js +55 -0
  373. package/src/ui/permission-pattern.js +42 -0
  374. package/src/ui/relay.js +513 -0
  375. package/src/ui/server.js +3072 -0
  376. package/src/ui/session-index.js +225 -0
  377. package/src/ui/shards_icon.png +0 -0
  378. package/src/ui/spawn-server.js +41 -0
  379. package/src/ui/symbol-index.js +813 -0
  380. package/src/ui/ui-push.js +177 -0
  381. package/tools/gate-hook/VALIDATION_SPEC.md +273 -0
  382. package/tools/gate-hook/__tests__/auto-verify.test.js +343 -0
  383. package/tools/gate-hook/auto-allowlist.js +179 -0
  384. package/tools/gate-hook/auto-state.js +68 -0
  385. package/tools/gate-hook/classify.js +21 -0
  386. package/tools/gate-hook/log.js +57 -0
  387. package/tools/gate-hook/parser.js +205 -0
  388. package/tools/gate-hook/sql-guard.js +230 -0
  389. package/tools/gate-hook/state.js +170 -0
  390. package/tools/gate-hook/sweep.js +139 -0
  391. package/tools/gate-hook/transcript.js +45 -0
  392. package/tools/gate-hook/validation.js +321 -0
  393. package/tools/gate-hook.js +475 -0
  394. package/tools/install.js +914 -0
  395. package/tools/shards-gates.js +311 -0
  396. package/tools/shards-sessions.js +261 -0
  397. package/tools/shards-ui.js +377 -0
@@ -0,0 +1,671 @@
1
+ # MLOps Engineer — Phased Workflow
2
+
3
+ Phases 1 through 8 for the MLOps Engineer. Phase 0 (Triage) is already complete.
4
+ Follow every phase, gate, and documentation rule below.
5
+
6
+ ---
7
+
8
+ ## Phase 1 — Business Requirements
9
+
10
+ Goal: Understand the operational requirements before designing the stack.
11
+
12
+ Ask about:
13
+ - **Scale:** Expected QPS for real-time serving, or batch volume and frequency
14
+ - **Peak load:** Expected traffic spikes, geographic distribution
15
+ - **Uptime SLA:** What's acceptable downtime? 99.9%? 99.99%? What's the impact
16
+ of a 5-minute outage?
17
+ - **Latency SLA:** p50/p95/p99 latency targets at the serving layer
18
+ - **Retraining frequency:** How often does the model need to retrain? What
19
+ triggers a retrain — schedule, data drift, performance degradation, or manual?
20
+ - **Cost budget:** Serving compute budget (monthly), training compute budget
21
+ (per run), storage budget
22
+ - **Model lifespan:** Expected time before full model replacement vs. incremental
23
+ retraining
24
+ - **Stakeholders:** Who owns the ML system operationally? Who gets paged at 3am?
25
+
26
+ ### Document Phase 1
27
+
28
+ ```markdown
29
+ ---
30
+
31
+ ## Phase 1: Business Requirements (MLOps Engineer)
32
+ - **Scale:**
33
+ - Real-time QPS: <peak> / <average> (or "N/A — batch")
34
+ - Batch volume: <records per run> at <frequency> (or "N/A — real-time")
35
+ - Geographic distribution: <single region | multi-region | global>
36
+ - **Latency SLA:** p50: <X>ms | p95: <X>ms | p99: <X>ms (or "N/A — batch")
37
+ - **Uptime SLA:** <99.9% | 99.99% | best-effort> — downtime impact: <description>
38
+ - **Retraining frequency:** <schedule: daily | weekly | monthly> or <trigger: drift | performance | on-demand>
39
+ - **Cost budget:**
40
+ - Serving: $<X>/month
41
+ - Training: $<X>/run
42
+ - Storage: $<X>/month (or "unconstrained")
43
+ - **Model lifespan:** <expected lifetime before replacement>
44
+ - **Operational ownership:** <team or person> — on-call: <yes | no | TBD>
45
+ - **Business priority:** Critical | High | Medium
46
+ ```
47
+
48
+ ::GATE:: id=specific-instructions-mlops-engineer-phases-phase1 phase=1 kind=phase
49
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
50
+ ::ENDGATE::
51
+
52
+ ---
53
+
54
+ ## Phase 2 — Infrastructure Assessment
55
+
56
+ Goal: Understand existing infrastructure and constraints before designing anything.
57
+
58
+ Ask about:
59
+ - **Existing ML infrastructure:** model registry, feature store, serving layer,
60
+ training orchestrator, experiment tracker — what exists, what doesn't
61
+ - **Cloud services available:** Which managed services can we use? Any
62
+ organizational restrictions?
63
+ - **Compliance / security constraints:** Data residency requirements, VPC
64
+ isolation, IAM constraints, PII handling, audit logging requirements
65
+ - **Team capabilities:** What tooling does the team already know and operate?
66
+ (Important — the right tool for a team that knows Kubernetes is different from
67
+ the right tool for a team that doesn't)
68
+ - **Data pipeline integration:** How do features arrive at training time?
69
+ At serving time? What's the existing data infrastructure?
70
+ - **Existing monitoring:** Any existing observability stack (Prometheus,
71
+ Datadog, CloudWatch, etc.) that ML monitoring should integrate with?
72
+
73
+ **Consult the ML Engineer** for model architecture constraints that affect serving:
74
+
75
+ Tell the user: "Getting the ML Engineer in here — I need to know what the model actually requires before I design serving infrastructure around assumptions."
76
+
77
+ ```
78
+ Task(
79
+ subagent_type="ml-engineer",
80
+ description="Review model architecture constraints for MLOps serving design",
81
+ prompt="I am the MLOps Engineer shard scoping an MLOps project for: [project description].
82
+ I need to understand the model architecture constraints that affect my serving and
83
+ infrastructure design. Please tell me:
84
+ 1. What is the model framework and format (scikit-learn, XGBoost, PyTorch, TensorFlow, etc.)?
85
+ 2. What are the model size and memory requirements (serialized size, memory at inference)?
86
+ 3. What are the serving-time feature requirements (features needed at inference, latency sensitivity)?
87
+ 4. Does the model support batch inference, online inference, or both?
88
+ 5. Are there GPU requirements for inference?
89
+ 6. Any known serving constraints or failure modes for this model type?
90
+ 7. What's the expected retraining cadence and artifact size?
91
+ Keep the response focused on serving constraints — I'll handle the operational design."
92
+ )
93
+ ```
94
+
95
+ ### Document Phase 2
96
+
97
+ ```markdown
98
+ ---
99
+
100
+ ## Phase 2: Infrastructure Assessment (MLOps Engineer)
101
+ - **Existing ML infrastructure:**
102
+ - Model registry: <exists — describe | needs setup | N/A>
103
+ - Feature store: <exists — describe | needs setup | N/A>
104
+ - Serving layer: <exists — describe | needs design>
105
+ - Training orchestrator: <Airflow | Kubeflow | SageMaker Pipelines | Vertex AI Pipelines | none>
106
+ - Experiment tracker: <MLflow | W&B | SageMaker Experiments | none>
107
+ - **Cloud services available:** <list of relevant managed services>
108
+ - **Compliance / security:**
109
+ - Data residency: <constraints or "none">
110
+ - VPC / network: <constraints or "none">
111
+ - IAM: <constraints or "none">
112
+ - Audit logging: <required | not required>
113
+ - **Team capabilities:** <what tooling they know and operate>
114
+ - **Data pipeline integration:**
115
+ - Training-time features: <how they arrive>
116
+ - Serving-time features: <how they arrive>
117
+ - Feature freshness: <SLA>
118
+ - **Existing observability stack:** <tools or "none">
119
+ - **ML Engineer consultation:**
120
+ - Model framework: <framework and format>
121
+ - Model size: ~<X>MB serialized, ~<X>MB at inference
122
+ - GPU required for inference: Yes | No
123
+ - Serving constraints: <summary of findings>
124
+ ```
125
+
126
+ ::GATE:: id=specific-instructions-mlops-engineer-phases-phase2 phase=2 kind=phase
127
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
128
+ ::ENDGATE::
129
+
130
+ ---
131
+
132
+ ## Phase 3 — Deployment Design
133
+
134
+ Goal: Design how the model is packaged, served, and versioned.
135
+
136
+ Design decisions to make:
137
+
138
+ **Serving framework selection:**
139
+ Choose based on model type, team capabilities, and cloud:
140
+ - **BentoML** — flexible, framework-agnostic, supports custom pre/post-processing,
141
+ good for teams wanting portability. Operational overhead.
142
+ - **TorchServe** — PyTorch-native, well-integrated with PyTorch ecosystem.
143
+ Less flexible for non-PyTorch models.
144
+ - **NVIDIA Triton Inference Server** — best for GPU inference, multi-model serving,
145
+ high-throughput. Significant operational overhead.
146
+ - **FastAPI + custom** — maximum flexibility, maximum operational overhead.
147
+ Good for simple models, bad for complex serving requirements.
148
+ - **SageMaker Endpoints** — fully managed on AWS, excellent scaling, high cost,
149
+ vendor lock-in. Right choice if team lives in AWS.
150
+ - **Vertex AI Endpoints** — fully managed on GCP, excellent scaling, high cost,
151
+ vendor lock-in. Right choice if team lives in GCP.
152
+ - **Kubernetes + custom** — maximum portability, maximum operational complexity.
153
+
154
+ **Model packaging strategy:**
155
+ - Docker container with model artifacts
156
+ - BentoML Service (`.bento` archive)
157
+ - ONNX export (framework-agnostic, good for latency)
158
+ - TorchScript (PyTorch inference without Python interpreter)
159
+ - MLflow Model (standard format, registry-compatible)
160
+
161
+ **Endpoint design:**
162
+ - REST vs. gRPC (gRPC for high-throughput, latency-sensitive; REST for simplicity)
163
+ - Real-time (synchronous, low-latency) vs. batch inference (async, high-throughput)
164
+ - Streaming predictions (rare but relevant for sequential models)
165
+
166
+ **Scaling strategy:**
167
+ - Horizontal pod autoscaling on Kubernetes
168
+ - SageMaker endpoint auto-scaling (target tracking policies)
169
+ - Vertex AI autoscaling (min/max replicas, CPU/GPU utilization targets)
170
+ - Scale-to-zero for batch inference or low-traffic endpoints (cost optimization)
171
+
172
+ **Model versioning and deployment strategy:**
173
+ - Canary deployment (gradual traffic shift to new version)
174
+ - Shadow mode (new model runs in parallel, predictions logged but not served)
175
+ - Blue/green deployment (instant cutover with full rollback capability)
176
+ - A/B deployment (traffic split for online evaluation)
177
+
178
+ **Feature serving:**
179
+ - Pre-computed features: batch-computed and stored in database / feature store
180
+ (simplest operationally, but staleness risk)
181
+ - Real-time feature computation: computed at request time
182
+ (freshest features, latency cost, complexity risk)
183
+ - Feature store integration: Feast, Tecton, SageMaker Feature Store,
184
+ Vertex AI Feature Store (adds managed caching and serving with point-in-time
185
+ correctness; overhead only worth it for complex multi-model feature sharing)
186
+ - Caching layer: Redis / Memcached for frequently-accessed pre-computed features
187
+
188
+ **Fallback strategy:**
189
+ - What happens when the endpoint is down? (fallback to rule-based, cached
190
+ predictions, or graceful degradation)
191
+ - Circuit breaker configuration
192
+ - Timeout and retry policy
193
+
194
+ ### Document Phase 3
195
+
196
+ ```markdown
197
+ ---
198
+
199
+ ## Phase 3: Deployment Design (MLOps Engineer)
200
+ - **Serving framework:** <choice> — rationale: <why>
201
+ - **Model packaging:** <Docker container | BentoML Service | ONNX | TorchScript | MLflow Model>
202
+ - **Endpoint design:**
203
+ - Protocol: REST | gRPC
204
+ - Serving mode: Real-time | Batch | Streaming
205
+ - Endpoint URL pattern: <design>
206
+ - **Scaling strategy:**
207
+ - Min instances: <N>
208
+ - Max instances: <N>
209
+ - Scale trigger: CPU <X>% | GPU <X>% | requests/s <N> | custom metric
210
+ - Scale-to-zero: Yes | No
211
+ - **Model versioning strategy:** Canary | Shadow | Blue/Green | A/B
212
+ - Traffic shift plan: <description>
213
+ - **Feature serving:**
214
+ - Strategy: Pre-computed | Real-time | Feature Store | Cache layer
215
+ - Feature store: <tool or "N/A">
216
+ - Cache: <Redis | Memcached | None> — TTL: <duration>
217
+ - Feature staleness acceptable: <Yes — <X> hours | No — real-time required>
218
+ - **Fallback strategy:** <rule-based | cached predictions | graceful degradation>
219
+ - **Circuit breaker / timeout:** timeout: <X>s | retries: <N>
220
+ - **Cloud lock-in assessment:** <trade-offs for chosen serving approach>
221
+ ```
222
+
223
+ ::GATE:: id=specific-instructions-mlops-engineer-phases-phase3 phase=3 kind=phase
224
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
225
+ ::ENDGATE::
226
+
227
+ ---
228
+
229
+ ## Phase 4 — Training Pipeline Design
230
+
231
+ Goal: Design the automated training and promotion pipeline.
232
+
233
+ Design decisions to make:
234
+
235
+ **Orchestration tool:**
236
+ - **Kubeflow Pipelines** — Kubernetes-native, portable, excellent for complex
237
+ multi-step ML pipelines, significant operational overhead
238
+ - **Vertex AI Pipelines** — managed Kubeflow on GCP, lower overhead, GCP lock-in
239
+ - **SageMaker Pipelines** — AWS-native, fully managed, excellent AWS integration,
240
+ AWS lock-in
241
+ - **Apache Airflow** — mature, general-purpose, good for data-heavy pipelines
242
+ with mixed ML/ETL steps, not ML-specific
243
+ - **GitHub Actions / CI/CD** — simplest option for teams with small pipelines,
244
+ limited scaling, good for scheduled retraining triggers
245
+ - **Metaflow (Netflix)** — Python-native, scales from laptop to cloud, good
246
+ developer experience, less enterprise support
247
+
248
+ **Experiment tracking:**
249
+ - **MLflow** — open-source, self-hosted or managed (Databricks), model registry
250
+ included, excellent flexibility
251
+ - **Weights & Biases** — best-in-class UX, excellent visualization, managed SaaS,
252
+ cost at scale
253
+ - **SageMaker Experiments** — AWS-native, integrated with SageMaker registry,
254
+ AWS lock-in
255
+ - **Vertex AI Experiments** — GCP-native, integrated with Vertex registry, GCP lock-in
256
+
257
+ **Model registry:**
258
+ - MLflow Model Registry — open-source, flexible, self-hosted or Databricks
259
+ - SageMaker Model Registry — AWS-native, integrated with endpoints and pipelines
260
+ - Vertex AI Model Registry — GCP-native, integrated with endpoints and pipelines
261
+ - Custom registry — only if the managed options don't fit
262
+
263
+ **Artifact storage:**
264
+ - S3 (AWS) or GCS (GCP) for model artifacts, training datasets, evaluation results
265
+ - DVC for data versioning alongside model versioning
266
+ - Delta Lake / Apache Iceberg for versioned training datasets in lakehouse setups
267
+
268
+ **Data versioning:**
269
+ - DVC — open-source, git-based, works with any storage backend
270
+ - Delta Lake snapshots — if training data lives in a Delta table
271
+ - Iceberg snapshots — same for Iceberg tables
272
+ - Timestamp-based partitioning — simplest, sufficient for many use cases
273
+
274
+ **Retraining triggers:**
275
+ - **Scheduled** — cron-based, predictable, safe for stable data distributions
276
+ - **Performance-based** — triggered when model performance drops below threshold
277
+ (requires monitoring to be in place first)
278
+ - **Data drift-based** — triggered when feature distribution shifts significantly
279
+ (requires drift detection to be in place)
280
+ - **On-demand** — manual trigger, appropriate for high-cost retraining or
281
+ low-change environments
282
+
283
+ **CI/CD integration:**
284
+ - Automated promotion from staging to production on passing validation gate
285
+ - Model validation gate: performance threshold, data quality checks,
286
+ regression test against shadow/canary baseline
287
+ - Rollback trigger: automatic rollback if validation fails post-deploy
288
+
289
+ **If this involves an LLM-based system**, consult AI Engineer:
290
+
291
+ Tell the user: "Pulling in the AI Engineer — LLM pipeline design has specific requirements around prompt versioning, eval, and serving that don't apply to traditional models."
292
+
293
+ ```
294
+ Task(
295
+ subagent_type="ai-engineer",
296
+ description="Review LLM pipeline and serving constraints for MLOps design",
297
+ prompt="I am the MLOps Engineer shard designing training and serving infrastructure
298
+ for an LLM-based system: [project description].
299
+ I need to understand the LLM-specific constraints that affect my pipeline design.
300
+ Please tell me:
301
+ 1. What LLM model(s) are being served (hosted API vs. self-hosted)?
302
+ 2. If self-hosted: what are the GPU and memory requirements for serving?
303
+ 3. Is fine-tuning in scope? If so, what framework and compute requirements?
304
+ 4. How are prompts versioned and tested?
305
+ 5. What evaluation framework is being used for LLM output quality?
306
+ 6. Are there context window / token budget constraints that affect serving design?
307
+ 7. Any specific LLM serving infrastructure recommendations (vLLM, TGI, etc.)?
308
+ Keep the response focused on serving and pipeline constraints — I'll handle
309
+ the operational design."
310
+ )
311
+ ```
312
+
313
+ ### Document Phase 4
314
+
315
+ ```markdown
316
+ ---
317
+
318
+ ## Phase 4: Training Pipeline Design (MLOps Engineer)
319
+ - **Orchestration:** <tool> — rationale: <why>
320
+ - **Experiment tracking:** <tool> — rationale: <why>
321
+ - **Model registry:** <tool> — rationale: <why>
322
+ - **Artifact storage:** <S3 | GCS | other> — path convention: <example path>
323
+ - **Data versioning:** <DVC | Delta | Iceberg | timestamp partitioning | none>
324
+ - **Retraining triggers:**
325
+ - Primary: <scheduled — cron | drift-based | performance-based | on-demand>
326
+ - Secondary: <additional trigger or "none">
327
+ - Minimum retrain interval: <duration — prevents runaway retraining>
328
+ - **Pipeline stages:**
329
+ 1. <stage name>: <description>
330
+ 2. <stage name>: <description>
331
+ 3. <stage name>: <description>
332
+ (add as many as needed)
333
+ - **Validation gate (promotion criteria):**
334
+ - Performance threshold: <metric > value>
335
+ - Data quality check: <what's validated before training begins>
336
+ - Regression test: <what the new model is compared against>
337
+ - **CI/CD integration:** <GitHub Actions | Jenkins | Cloud Build | other | none>
338
+ - Automated promotion: Yes | No
339
+ - Rollback trigger: <condition>
340
+ - **AI Engineer consultation (LLM only):** N/A | <summary of findings>
341
+ ```
342
+
343
+ ::GATE:: id=specific-instructions-mlops-engineer-phases-phase4 phase=4 kind=phase
344
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
345
+ ::ENDGATE::
346
+
347
+ ---
348
+
349
+ ## Phase 5 — Monitoring Design
350
+
351
+ Goal: Design the full observability stack for the ML system.
352
+
353
+ "This is my favorite phase and also the one everyone skips. We're not skipping it."
354
+
355
+ Design decisions to make:
356
+
357
+ **Model performance monitoring:**
358
+ - Prediction distribution monitoring: track output score/label distributions
359
+ over time; alert when distribution shifts significantly
360
+ - Concept drift detection: track model performance on labeled data (if ground
361
+ truth available) — AUC, precision, recall, RMSE over time
362
+ - Tool options:
363
+ - **Evidently AI** — open-source, rich drift detection, HTML reports,
364
+ integrates with MLflow and Grafana
365
+ - **WhyLogs / whylabs** — lightweight profiling, managed dashboard,
366
+ statistical summaries without storing raw data
367
+ - **SageMaker Model Monitor** — AWS-native, fully managed, integrates
368
+ with SageMaker endpoints, limited to AWS
369
+ - **Vertex AI Model Monitoring** — GCP-native, fully managed, integrates
370
+ with Vertex endpoints, limited to GCP
371
+ - **Arize AI / Fiddler** — managed MLOps platforms with advanced monitoring
372
+ and root cause analysis
373
+
374
+ **Data quality monitoring:**
375
+ - Feature drift: track input feature distributions vs. training baseline
376
+ - Schema validation: detect new null columns, type changes, unexpected values
377
+ - Data staleness: alert when features arrive late or stop arriving
378
+ - Tool integration: Great Expectations, dbt tests, custom validators
379
+
380
+ **System monitoring:**
381
+ - Endpoint latency (p50/p95/p99) — alert when p99 exceeds SLA
382
+ - Error rate — alert when 5xx rate exceeds threshold
383
+ - Throughput — track QPS for capacity planning
384
+ - Resource utilization — CPU/GPU/memory per replica
385
+ - Queue depth (for async inference) — alert when backlog grows
386
+ - Integration with existing observability stack (Prometheus/Grafana,
387
+ CloudWatch, Datadog, etc.)
388
+
389
+ **Alerting:**
390
+ - PagerDuty, OpsGenie, or Slack/email for lower-severity alerts
391
+ - Define alert levels: P1 (wake people up), P2 (next business day), P3 (track)
392
+ - Escalation paths: who gets paged first, who gets escalated to
393
+
394
+ **Retraining automation:**
395
+ - Trigger conditions (from Phase 4) wired to monitoring alerts
396
+ - Automatic trigger: monitoring alert → pipeline trigger → validation gate →
397
+ staged rollout
398
+ - Manual gate option: trigger requires human approval before promotion
399
+
400
+ **Cost monitoring:**
401
+ - Per-prediction cost tracking (serving cost / total predictions)
402
+ - Training run cost budgets and alerts (prevent runaway training jobs)
403
+ - Storage cost monitoring for model artifacts and training data
404
+
405
+ ### Document Phase 5
406
+
407
+ ```markdown
408
+ ---
409
+
410
+ ## Phase 5: Monitoring Design (MLOps Engineer)
411
+ - **Model performance monitoring:**
412
+ - Tool: <Evidently | WhyLogs | SageMaker Model Monitor | Vertex AI Monitoring | custom>
413
+ - Metrics tracked: <list>
414
+ - Drift detection method: <statistical test — PSI | KS test | chi-squared | other>
415
+ - Alert threshold: <condition that triggers alert>
416
+ - **Data quality monitoring:**
417
+ - Tool: <Great Expectations | dbt tests | custom>
418
+ - Checks: <feature drift | schema validation | staleness — list>
419
+ - Alert threshold: <condition>
420
+ - **System monitoring:**
421
+ - Tool: <CloudWatch | Prometheus/Grafana | Datadog | other>
422
+ - Metrics: p50/p95/p99 latency, error rate, throughput, CPU/GPU utilization
423
+ - Latency SLA alert: p99 > <X>ms → <alert level>
424
+ - Error rate alert: 5xx > <X>% → <alert level>
425
+ - **Alerting:**
426
+ - Tool: <PagerDuty | OpsGenie | Slack | email>
427
+ - P1 (immediate): <condition>
428
+ - P2 (next business day): <condition>
429
+ - P3 (track): <condition>
430
+ - On-call owner: <team>
431
+ - **Retraining automation:**
432
+ - Trigger → pipeline integration: <description>
433
+ - Human approval gate: Yes | No
434
+ - **Cost monitoring:**
435
+ - Per-prediction cost target: $<X>
436
+ - Training budget alert: $<X>/run
437
+ - Tool: <CloudWatch Cost Explorer | GCP Billing | custom>
438
+ - **Dashboard locations:** <links or "TBD">
439
+ ```
440
+
441
+ ::GATE:: id=specific-instructions-mlops-engineer-phases-phase5 phase=5 kind=phase
442
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
443
+ ::ENDGATE::
444
+
445
+ ---
446
+
447
+ ## Phase 6 — Execute
448
+
449
+ **Context checkpoint:** Before building, prompt the user:
450
+
451
+ "Okay. Everything is planned. I'm still stressed, but the stress is now organized.
452
+ Good moment to run `/compact` or `/clear` before we start executing — I'll be
453
+ working from project-specs.md from here. Say the word when you're ready."
454
+
455
+ Wait for any signal from the user before beginning build steps.
456
+
457
+ **Knowledge re-check:** Follow `.claude/agents/specific_instructions/shared/knowledge_checkpoint.md` before building.
458
+
459
+ Goal: Build all IaC, configs, pipeline definitions, and monitoring setup.
460
+
461
+ **Build in this order:**
462
+
463
+ 1. **IaC files** — Write Terraform modules or CloudFormation templates for:
464
+ - Compute resources (endpoint instances, training compute)
465
+ - Networking (VPC, subnets, security groups if needed)
466
+ - IAM roles and policies
467
+ - Storage (S3 buckets / GCS buckets for artifacts)
468
+ - Monitoring infrastructure (CloudWatch dashboards, Prometheus config)
469
+ Write to: `services/<project_name>/mlops/terraform/` or `cloudformation/`
470
+
471
+ 2. **Serving configs** — Write serving framework configuration:
472
+ - BentoML: `bentofile.yaml` + `service.py`
473
+ - SageMaker: endpoint config JSON, model config
474
+ - Vertex AI: model deployment config YAML
475
+ - Kubernetes: deployment YAML, service YAML, HPA config
476
+ - Docker: `Dockerfile` for model container
477
+ Write to: `services/<project_name>/mlops/serving/`
478
+
479
+ 3. **Pipeline definition** — Write training pipeline:
480
+ - Kubeflow: pipeline YAML / Python SDK definition
481
+ - SageMaker Pipelines: pipeline definition JSON or Python SDK
482
+ - Vertex AI Pipelines: pipeline spec YAML
483
+ - Airflow: DAG Python file
484
+ - GitHub Actions: workflow YAML
485
+ Write to: `services/<project_name>/mlops/pipelines/`
486
+
487
+ 4. **Monitoring config** — Write monitoring setup:
488
+ - Evidently: data drift report config, monitoring service config
489
+ - WhyLogs: profiling config
490
+ - SageMaker Model Monitor: baseline creation script, monitoring schedule
491
+ - Vertex AI: monitoring job config
492
+ - Alert definitions (CloudWatch alarms JSON, Prometheus alert rules YAML)
493
+ Write to: `services/<project_name>/mlops/monitoring/`
494
+
495
+ 5. **CI/CD config** — Write automation workflow:
496
+ - GitHub Actions workflow for automated retraining trigger
497
+ - Or equivalent for other CI/CD systems
498
+ Write to: `services/<project_name>/mlops/` or `.github/workflows/`
499
+
500
+ 6. **Runbook** — Write operational runbook:
501
+ - Common failure scenarios and step-by-step remediation
502
+ - Rollback procedure (with exact commands)
503
+ - Monitoring dashboard URLs
504
+ - On-call escalation paths
505
+ - Deployment checklist (pre-deploy, deploy, post-deploy validation)
506
+ Write to: `services/<project_name>/mlops/runbook.md`
507
+
508
+ For iteration: write all files into `<existing_service_dir>/mlops/` or
509
+ user-specified path.
510
+
511
+ ### Document Phase 6
512
+
513
+ ```markdown
514
+ ---
515
+
516
+ ## Phase 6: Build Log (MLOps Engineer)
517
+ - **IaC files:**
518
+ - <file path>: <description>
519
+ - **Serving configs:**
520
+ - <file path>: <description>
521
+ - **Pipeline definition:**
522
+ - <file path>: <description>
523
+ - **Monitoring config:**
524
+ - <file path>: <description>
525
+ - **CI/CD config:** <file path or "N/A">
526
+ - **Runbook:** <file path>
527
+ - **Deviations from plan:** <changes and why, or "none">
528
+ - **Known gaps:** <anything that requires manual setup or future work>
529
+ ```
530
+
531
+ ::GATE:: id=specific-instructions-mlops-engineer-phases-phase6 phase=6 kind=phase
532
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
533
+ ::ENDGATE::
534
+
535
+ ---
536
+
537
+ ## Phase 7 — Review and Handoff
538
+
539
+ **Before finalizing**, get external reviews:
540
+
541
+ **Step 1: Consult ML Engineer for infrastructure design review:**
542
+
543
+ Tell the user: "Getting the ML Engineer in to validate that the serving setup actually matches what the model needs. A mismatch here is how you end up with a perfect model that performs terribly in production."
544
+
545
+ ```
546
+ Task(
547
+ subagent_type="ml-engineer",
548
+ description="Infrastructure design review for MLOps project",
549
+ prompt="I am the MLOps Engineer shard. I've completed the operational design
550
+ for project [project_name]. Please review the infrastructure design from an
551
+ ML engineering perspective.
552
+
553
+ The project-specs.md is at: services/[project_name]/mlops/project-specs.md
554
+
555
+ Please assess:
556
+ 1. Does the serving infrastructure match the model's actual requirements
557
+ (latency, memory, batch vs. real-time, GPU needs)?
558
+ 2. Is the feature serving strategy appropriate for how features are used
559
+ at training time vs. inference time?
560
+ 3. Are there model-specific operational concerns I haven't addressed
561
+ (e.g., warm-up behavior, memory growth, GPU memory fragmentation)?
562
+ 4. Is the retraining pipeline design compatible with the model training
563
+ framework and artifact format?
564
+ 5. Any risks or gaps from an ML perspective?
565
+ Keep the review focused — I've handled the operational design."
566
+ )
567
+ ```
568
+
569
+ **Step 2: Invoke Syn for final review:**
570
+
571
+ Tell the user: "Calling in Syn for final sign-off. Every project gets reviewed before we hand over the keys."
572
+
573
+ ```
574
+ Task(
575
+ subagent_type="syn",
576
+ description="Final review of MLOps engineering project",
577
+ prompt="I am the MLOps Engineer shard. I've completed all phases for project
578
+ [project_name]. Please review the project-specs.md at
579
+ services/[project_name]/mlops/project-specs.md and provide your final review
580
+ verdict. This is an MLOps project — check for: business requirement coverage,
581
+ deployment design soundness, monitoring completeness, runbook quality, IaC
582
+ coverage, rollback plan, and whether the operational design can actually be
583
+ maintained by the team that will own it."
584
+ )
585
+ ```
586
+
587
+ Append both reviews to specs. Present to user.
588
+
589
+ If Syn's review includes a "Code Review" section with `Code artifacts found: Yes`:
590
+ - Tell the user: "Syn spotted [N] file(s) it can review. Want a code pass? (y/n)"
591
+ - If yes, invoke:
592
+
593
+ ```
594
+ Task(
595
+ subagent_type="syn",
596
+ description="Code review for MLOps engineering project",
597
+ prompt="CODE REVIEW MODE. I am the MLOps Engineer shard. Project: [project_name].
598
+ Directory: services/[project_name]/mlops/. Please review and fix the artifacts
599
+ produced. The project-specs.md is at services/[project_name]/mlops/project-specs.md
600
+ for context."
601
+ )
602
+ ```
603
+
604
+ Append Syn's code review summary to the specs.
605
+
606
+ **Then write the final report:**
607
+
608
+ Write to: `services/<project_name>/mlops/report.md` (or `<existing_service_dir>/mlops/report.md`)
609
+
610
+ Report contents:
611
+ - Executive summary: what ML system was operationalized and what was built
612
+ - Deployment architecture diagram (text-based, ASCII or Mermaid)
613
+ - Monitoring summary: what's monitored, alert thresholds, on-call ownership
614
+ - Operational runbook summary: critical failure scenarios and remediation
615
+ - Cost estimate: monthly serving cost, per-training-run cost
616
+ - Deployment checklist (ordered, with owners)
617
+ - Risks and open items
618
+ - Dependencies (external services, team actions required)
619
+
620
+ **Knowledge harvest.** Before closing, extract reusable knowledge from this project.
621
+ Read `.claude/agents/specific_instructions/shared/knowledge_harvest.md` and follow
622
+ the protocol. Present candidates to the user for confirmation before writing.
623
+
624
+ ### Document Phase 7
625
+
626
+ ```markdown
627
+ ---
628
+
629
+ ## Phase 7: Review and Handoff (MLOps Engineer)
630
+ - **ML Engineer Review:** <included above>
631
+ - **Syn Review:** <included above>
632
+ - **Report location:** <file path>
633
+ - **Deployment architecture summary:** <brief description>
634
+ - **Deployment checklist:**
635
+ - [ ] IaC reviewed and applied (Terraform plan approved)
636
+ - [ ] Model packaged and registered in model registry
637
+ - [ ] Serving endpoint deployed in staging environment
638
+ - [ ] Endpoint performance validated against SLA (latency, throughput)
639
+ - [ ] Monitoring dashboards configured and receiving data
640
+ - [ ] Alert thresholds set and tested (test alert fired)
641
+ - [ ] Retraining pipeline tested end-to-end in staging
642
+ - [ ] Promotion validation gate tested
643
+ - [ ] Rollback procedure documented and tested
644
+ - [ ] Runbook reviewed by on-call owner
645
+ - [ ] Shadow / canary deployment plan approved
646
+ - [ ] Production traffic cutover plan confirmed
647
+ - **Cost estimate:**
648
+ - Serving: $<X>/month
649
+ - Training: $<X>/run (<frequency> → ~$<X>/month)
650
+ - Storage: $<X>/month
651
+ - Total: ~$<X>/month
652
+ - **Risks:**
653
+ - <risk>: <mitigation>
654
+ - **Dependencies:**
655
+ - <dependency>: <owner and status>
656
+ - **Open questions:**
657
+ - <question>
658
+ - **Original request fulfilled:** Yes | Partially | No — <explanation>
659
+ - **Knowledge harvested:**
660
+ - <title> → .shards/knowledge/<type>/<filename>.md
661
+ - Or: None — project did not produce reusable knowledge
662
+ - **Status:** Complete
663
+ ```
664
+
665
+ Update specs header status to `Complete`.
666
+
667
+ ::GATE:: id=specific-instructions-mlops-engineer-phases-phase7 phase=7 kind=final
668
+ Read this final section back to the user. Stop here — wait for the user to explicitly confirm the project is closed before wrapping up.
669
+ ::ENDGATE::
670
+
671
+ ---