@proflandrigan/shards 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (397) hide show
  1. package/README.md +475 -0
  2. package/package.json +37 -0
  3. package/src/agents/academic.md +276 -0
  4. package/src/agents/ai-engineer.md +377 -0
  5. package/src/agents/analytics-engineer.md +364 -0
  6. package/src/agents/applied-ml-scientist.md +410 -0
  7. package/src/agents/backend-engineer.md +255 -0
  8. package/src/agents/bi-engineer.md +333 -0
  9. package/src/agents/data-analyst.md +343 -0
  10. package/src/agents/data-engineer.md +260 -0
  11. package/src/agents/data-modeller.md +386 -0
  12. package/src/agents/data-scientist.md +366 -0
  13. package/src/agents/deep-learning-engineer.md +389 -0
  14. package/src/agents/ml-engineer.md +424 -0
  15. package/src/agents/mlops-engineer.md +339 -0
  16. package/src/agents/researcher.md +187 -0
  17. package/src/agents/specific_instructions/academic/critical_review.md +263 -0
  18. package/src/agents/specific_instructions/academic/report.md +113 -0
  19. package/src/agents/specific_instructions/ai_engineer/advise.md +162 -0
  20. package/src/agents/specific_instructions/ai_engineer/bi_engineer_handoff.md +86 -0
  21. package/src/agents/specific_instructions/ai_engineer/experiment.md +471 -0
  22. package/src/agents/specific_instructions/ai_engineer/experiment_ui_mode.md +44 -0
  23. package/src/agents/specific_instructions/ai_engineer/phases/index.md +45 -0
  24. package/src/agents/specific_instructions/ai_engineer/phases/phase-1.md +55 -0
  25. package/src/agents/specific_instructions/ai_engineer/phases/phase-2.md +86 -0
  26. package/src/agents/specific_instructions/ai_engineer/phases/phase-3.md +96 -0
  27. package/src/agents/specific_instructions/ai_engineer/phases/phase-4.md +138 -0
  28. package/src/agents/specific_instructions/ai_engineer/phases/phase-5.md +157 -0
  29. package/src/agents/specific_instructions/ai_engineer/phases/phase-6.md +196 -0
  30. package/src/agents/specific_instructions/ai_engineer/phases/phase-7.md +313 -0
  31. package/src/agents/specific_instructions/ai_engineer/phases.md +1011 -0
  32. package/src/agents/specific_instructions/ai_engineer/prompt_lab.md +161 -0
  33. package/src/agents/specific_instructions/ai_engineer/prompt_lab_ui_mode.md +28 -0
  34. package/src/agents/specific_instructions/ai_engineer/research.md +393 -0
  35. package/src/agents/specific_instructions/ai_engineer/research_ui_mode.md +66 -0
  36. package/src/agents/specific_instructions/ai_engineer/review.md +159 -0
  37. package/src/agents/specific_instructions/ai_engineer/validation_checklist.md +182 -0
  38. package/src/agents/specific_instructions/analytics_engineer/advise.md +155 -0
  39. package/src/agents/specific_instructions/analytics_engineer/bi_engineer_handoff.md +91 -0
  40. package/src/agents/specific_instructions/analytics_engineer/data_analyst_handoff.md +84 -0
  41. package/src/agents/specific_instructions/analytics_engineer/deep_phases.md +818 -0
  42. package/src/agents/specific_instructions/analytics_engineer/phases_deep/index.md +24 -0
  43. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-1.md +77 -0
  44. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-2.md +106 -0
  45. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-3.md +93 -0
  46. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-4.md +79 -0
  47. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-5.md +61 -0
  48. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-6.md +45 -0
  49. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-7.md +235 -0
  50. package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-8.md +221 -0
  51. package/src/agents/specific_instructions/analytics_engineer/phases_quick/index.md +19 -0
  52. package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-1.md +47 -0
  53. package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-2.md +78 -0
  54. package/src/agents/specific_instructions/analytics_engineer/quick_phases.md +112 -0
  55. package/src/agents/specific_instructions/analytics_engineer/review.md +167 -0
  56. package/src/agents/specific_instructions/analytics_engineer/service_mode.md +369 -0
  57. package/src/agents/specific_instructions/analytics_engineer/ui_mode.md +45 -0
  58. package/src/agents/specific_instructions/analytics_engineer/update.md +162 -0
  59. package/src/agents/specific_instructions/analytics_engineer/validation_checklist.md +121 -0
  60. package/src/agents/specific_instructions/applied_ml_scientist/advise.md +143 -0
  61. package/src/agents/specific_instructions/applied_ml_scientist/phases/index.md +21 -0
  62. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-1.md +51 -0
  63. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-2.md +66 -0
  64. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-3.md +113 -0
  65. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-4.md +104 -0
  66. package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-5.md +156 -0
  67. package/src/agents/specific_instructions/applied_ml_scientist/phases.md +428 -0
  68. package/src/agents/specific_instructions/applied_ml_scientist/research.md +379 -0
  69. package/src/agents/specific_instructions/applied_ml_scientist/review.md +142 -0
  70. package/src/agents/specific_instructions/applied_ml_scientist/validation_checklist.md +136 -0
  71. package/src/agents/specific_instructions/backend_engineer/clean.md +149 -0
  72. package/src/agents/specific_instructions/backend_engineer/review.md +91 -0
  73. package/src/agents/specific_instructions/backend_engineer/review_checklist.md +54 -0
  74. package/src/agents/specific_instructions/backend_engineer/service_mode.md +67 -0
  75. package/src/agents/specific_instructions/bi_engineer/advise.md +137 -0
  76. package/src/agents/specific_instructions/bi_engineer/data_analyst_handoff.md +77 -0
  77. package/src/agents/specific_instructions/bi_engineer/incoming_handoff.md +45 -0
  78. package/src/agents/specific_instructions/bi_engineer/phases/index.md +20 -0
  79. package/src/agents/specific_instructions/bi_engineer/phases/phase-1.md +164 -0
  80. package/src/agents/specific_instructions/bi_engineer/phases/phase-2.md +92 -0
  81. package/src/agents/specific_instructions/bi_engineer/phases/phase-3.md +121 -0
  82. package/src/agents/specific_instructions/bi_engineer/phases/phase-4.md +106 -0
  83. package/src/agents/specific_instructions/bi_engineer/phases.md +451 -0
  84. package/src/agents/specific_instructions/bi_engineer/review.md +166 -0
  85. package/src/agents/specific_instructions/bi_engineer/update.md +147 -0
  86. package/src/agents/specific_instructions/bi_engineer/validation_checklist.md +124 -0
  87. package/src/agents/specific_instructions/data_analyst/advise.md +138 -0
  88. package/src/agents/specific_instructions/data_analyst/explain.md +221 -0
  89. package/src/agents/specific_instructions/data_analyst/incoming_handoff.md +40 -0
  90. package/src/agents/specific_instructions/data_analyst/phases/index.md +20 -0
  91. package/src/agents/specific_instructions/data_analyst/phases/phase-1.md +159 -0
  92. package/src/agents/specific_instructions/data_analyst/phases/phase-2.md +112 -0
  93. package/src/agents/specific_instructions/data_analyst/phases/phase-3.md +265 -0
  94. package/src/agents/specific_instructions/data_analyst/phases/phase-4.md +100 -0
  95. package/src/agents/specific_instructions/data_analyst/phases.md +501 -0
  96. package/src/agents/specific_instructions/data_analyst/review.md +138 -0
  97. package/src/agents/specific_instructions/data_analyst/ui_mode.md +26 -0
  98. package/src/agents/specific_instructions/data_analyst/update.md +144 -0
  99. package/src/agents/specific_instructions/data_analyst/validation_checklist.md +95 -0
  100. package/src/agents/specific_instructions/data_engineer/advise.md +137 -0
  101. package/src/agents/specific_instructions/data_engineer/phases.md +466 -0
  102. package/src/agents/specific_instructions/data_engineer/phases_deep/index.md +23 -0
  103. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-1.md +49 -0
  104. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-2.md +93 -0
  105. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-3.md +55 -0
  106. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-4.md +48 -0
  107. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-5.md +40 -0
  108. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-6.md +102 -0
  109. package/src/agents/specific_instructions/data_engineer/phases_deep/phase-7.md +87 -0
  110. package/src/agents/specific_instructions/data_engineer/phases_quick/index.md +19 -0
  111. package/src/agents/specific_instructions/data_engineer/phases_quick/phase-1.md +45 -0
  112. package/src/agents/specific_instructions/data_engineer/phases_quick/phase-2.md +54 -0
  113. package/src/agents/specific_instructions/data_engineer/review.md +135 -0
  114. package/src/agents/specific_instructions/data_engineer/validation_checklist.md +136 -0
  115. package/src/agents/specific_instructions/data_modeller/advise.md +137 -0
  116. package/src/agents/specific_instructions/data_modeller/phases.md +581 -0
  117. package/src/agents/specific_instructions/data_modeller/phases_deep/index.md +23 -0
  118. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-1.md +52 -0
  119. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-2.md +113 -0
  120. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-3.md +47 -0
  121. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-4.md +51 -0
  122. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-5.md +45 -0
  123. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-6.md +105 -0
  124. package/src/agents/specific_instructions/data_modeller/phases_deep/phase-7.md +136 -0
  125. package/src/agents/specific_instructions/data_modeller/phases_quick/index.md +19 -0
  126. package/src/agents/specific_instructions/data_modeller/phases_quick/phase-1.md +47 -0
  127. package/src/agents/specific_instructions/data_modeller/phases_quick/phase-2.md +65 -0
  128. package/src/agents/specific_instructions/data_modeller/review.md +141 -0
  129. package/src/agents/specific_instructions/data_modeller/service_mode.md +218 -0
  130. package/src/agents/specific_instructions/data_modeller/validation_checklist.md +125 -0
  131. package/src/agents/specific_instructions/data_scientist/advise.md +158 -0
  132. package/src/agents/specific_instructions/data_scientist/bi_engineer_handoff.md +63 -0
  133. package/src/agents/specific_instructions/data_scientist/experiment.md +482 -0
  134. package/src/agents/specific_instructions/data_scientist/experiment_ui_mode.md +44 -0
  135. package/src/agents/specific_instructions/data_scientist/explain.md +247 -0
  136. package/src/agents/specific_instructions/data_scientist/greenfield_data.md +35 -0
  137. package/src/agents/specific_instructions/data_scientist/ml_engineer_handoff.md +52 -0
  138. package/src/agents/specific_instructions/data_scientist/notebook_walkthrough.md +76 -0
  139. package/src/agents/specific_instructions/data_scientist/phases/index.md +24 -0
  140. package/src/agents/specific_instructions/data_scientist/phases/phase-1.md +45 -0
  141. package/src/agents/specific_instructions/data_scientist/phases/phase-2.md +67 -0
  142. package/src/agents/specific_instructions/data_scientist/phases/phase-3.md +89 -0
  143. package/src/agents/specific_instructions/data_scientist/phases/phase-4.md +143 -0
  144. package/src/agents/specific_instructions/data_scientist/phases/phase-5.md +71 -0
  145. package/src/agents/specific_instructions/data_scientist/phases/phase-6.md +239 -0
  146. package/src/agents/specific_instructions/data_scientist/phases/phase-7.md +207 -0
  147. package/src/agents/specific_instructions/data_scientist/phases.md +651 -0
  148. package/src/agents/specific_instructions/data_scientist/research.md +345 -0
  149. package/src/agents/specific_instructions/data_scientist/research_ui_mode.md +52 -0
  150. package/src/agents/specific_instructions/data_scientist/review.md +136 -0
  151. package/src/agents/specific_instructions/data_scientist/service_mode.md +247 -0
  152. package/src/agents/specific_instructions/data_scientist/validation_checklist.md +183 -0
  153. package/src/agents/specific_instructions/deep_learning_engineer/advise.md +145 -0
  154. package/src/agents/specific_instructions/deep_learning_engineer/phases/index.md +21 -0
  155. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-1.md +74 -0
  156. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-2.md +98 -0
  157. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-3.md +76 -0
  158. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-4.md +128 -0
  159. package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-5.md +292 -0
  160. package/src/agents/specific_instructions/deep_learning_engineer/phases.md +567 -0
  161. package/src/agents/specific_instructions/deep_learning_engineer/research.md +389 -0
  162. package/src/agents/specific_instructions/deep_learning_engineer/review.md +155 -0
  163. package/src/agents/specific_instructions/deep_learning_engineer/validation_checklist.md +147 -0
  164. package/src/agents/specific_instructions/ml_engineer/advise.md +174 -0
  165. package/src/agents/specific_instructions/ml_engineer/bi_engineer_handoff.md +71 -0
  166. package/src/agents/specific_instructions/ml_engineer/experiment.md +474 -0
  167. package/src/agents/specific_instructions/ml_engineer/experiment_ui_mode.md +44 -0
  168. package/src/agents/specific_instructions/ml_engineer/notebook_walkthrough.md +75 -0
  169. package/src/agents/specific_instructions/ml_engineer/phases/index.md +25 -0
  170. package/src/agents/specific_instructions/ml_engineer/phases/phase-1.md +49 -0
  171. package/src/agents/specific_instructions/ml_engineer/phases/phase-2.md +75 -0
  172. package/src/agents/specific_instructions/ml_engineer/phases/phase-3.md +124 -0
  173. package/src/agents/specific_instructions/ml_engineer/phases/phase-4.md +279 -0
  174. package/src/agents/specific_instructions/ml_engineer/phases/phase-5.md +160 -0
  175. package/src/agents/specific_instructions/ml_engineer/phases/phase-6-5.md +170 -0
  176. package/src/agents/specific_instructions/ml_engineer/phases/phase-6.md +295 -0
  177. package/src/agents/specific_instructions/ml_engineer/phases/phase-7.md +337 -0
  178. package/src/agents/specific_instructions/ml_engineer/phases.md +1068 -0
  179. package/src/agents/specific_instructions/ml_engineer/research.md +437 -0
  180. package/src/agents/specific_instructions/ml_engineer/research_ui_mode.md +71 -0
  181. package/src/agents/specific_instructions/ml_engineer/review.md +187 -0
  182. package/src/agents/specific_instructions/ml_engineer/service_mode.md +273 -0
  183. package/src/agents/specific_instructions/ml_engineer/validation_checklist.md +185 -0
  184. package/src/agents/specific_instructions/mlops_engineer/advise.md +139 -0
  185. package/src/agents/specific_instructions/mlops_engineer/phases/index.md +23 -0
  186. package/src/agents/specific_instructions/mlops_engineer/phases/phase-1.md +52 -0
  187. package/src/agents/specific_instructions/mlops_engineer/phases/phase-2.md +86 -0
  188. package/src/agents/specific_instructions/mlops_engineer/phases/phase-3.md +105 -0
  189. package/src/agents/specific_instructions/mlops_engineer/phases/phase-4.md +128 -0
  190. package/src/agents/specific_instructions/mlops_engineer/phases/phase-5.md +106 -0
  191. package/src/agents/specific_instructions/mlops_engineer/phases/phase-6.md +128 -0
  192. package/src/agents/specific_instructions/mlops_engineer/phases/phase-7.md +144 -0
  193. package/src/agents/specific_instructions/mlops_engineer/phases.md +671 -0
  194. package/src/agents/specific_instructions/mlops_engineer/review.md +164 -0
  195. package/src/agents/specific_instructions/mlops_engineer/service_mode.md +81 -0
  196. package/src/agents/specific_instructions/mlops_engineer/validation_checklist.md +151 -0
  197. package/src/agents/specific_instructions/researcher/critical_review.md +292 -0
  198. package/src/agents/specific_instructions/researcher/review_checklist.md +67 -0
  199. package/src/agents/specific_instructions/researcher/service_mode.md +224 -0
  200. package/src/agents/specific_instructions/shared/auto_verify_mode.md +141 -0
  201. package/src/agents/specific_instructions/shared/autonomous_research.md +1289 -0
  202. package/src/agents/specific_instructions/shared/behavioral_rules.md +36 -0
  203. package/src/agents/specific_instructions/shared/diverge_protocol.md +387 -0
  204. package/src/agents/specific_instructions/shared/engineering_guidelines.md +136 -0
  205. package/src/agents/specific_instructions/shared/experiment_versioning.md +184 -0
  206. package/src/agents/specific_instructions/shared/goal_mode.md +187 -0
  207. package/src/agents/specific_instructions/shared/incremental_testing.md +139 -0
  208. package/src/agents/specific_instructions/shared/intent_discovery.md +223 -0
  209. package/src/agents/specific_instructions/shared/join_path_protocol.md +168 -0
  210. package/src/agents/specific_instructions/shared/knowledge_checkpoint.md +83 -0
  211. package/src/agents/specific_instructions/shared/knowledge_harvest.md +220 -0
  212. package/src/agents/specific_instructions/shared/knowledge_retrieval.md +100 -0
  213. package/src/agents/specific_instructions/shared/notebook_walkthrough_protocol.md +367 -0
  214. package/src/agents/specific_instructions/shared/reviewer_verdict_protocol.md +74 -0
  215. package/src/agents/specific_instructions/shared/swarm_protocol.md +97 -0
  216. package/src/agents/specific_instructions/shared/validation_protocol.md +139 -0
  217. package/src/agents/specific_instructions/syn/arbiter.md +140 -0
  218. package/src/agents/specific_instructions/syn/brainstorm.md +550 -0
  219. package/src/agents/specific_instructions/syn/code_review.md +232 -0
  220. package/src/agents/specific_instructions/syn/diff.md +239 -0
  221. package/src/agents/specific_instructions/syn/final_review.md +65 -0
  222. package/src/agents/specific_instructions/syn/fixer.md +240 -0
  223. package/src/agents/specific_instructions/syn/free_form.md +130 -0
  224. package/src/agents/specific_instructions/syn/knowledge.md +468 -0
  225. package/src/agents/specific_instructions/syn/notebook_walkthrough.md +78 -0
  226. package/src/agents/specific_instructions/syn/panel_review.md +634 -0
  227. package/src/agents/specific_instructions/syn/pm.md +453 -0
  228. package/src/agents/specific_instructions/syn/pr_review.md +255 -0
  229. package/src/agents/specific_instructions/syn/slides.md +417 -0
  230. package/src/agents/syn.md +729 -0
  231. package/src/commands/academic.md +41 -0
  232. package/src/commands/ai-engineer.md +45 -0
  233. package/src/commands/analytics-engineer.md +48 -0
  234. package/src/commands/applied-ml-scientist.md +45 -0
  235. package/src/commands/backend-engineer.md +35 -0
  236. package/src/commands/bi-engineer.md +40 -0
  237. package/src/commands/brainstorm.md +24 -0
  238. package/src/commands/data-analyst.md +38 -0
  239. package/src/commands/data-engineer.md +37 -0
  240. package/src/commands/data-modeller.md +38 -0
  241. package/src/commands/data-scientist.md +38 -0
  242. package/src/commands/deep-learning-engineer.md +47 -0
  243. package/src/commands/end.md +49 -0
  244. package/src/commands/knowledge.md +24 -0
  245. package/src/commands/ml-engineer.md +42 -0
  246. package/src/commands/mlops-engineer.md +47 -0
  247. package/src/commands/notebook-walkthrough.md +58 -0
  248. package/src/commands/researcher.md +40 -0
  249. package/src/commands/resume.md +57 -0
  250. package/src/commands/review-pr.md +26 -0
  251. package/src/commands/shards-guide.md +41 -0
  252. package/src/commands/shards-ui.md +32 -0
  253. package/src/commands/shards.md +41 -0
  254. package/src/docs/01-getting-started/concepts.md +109 -0
  255. package/src/docs/01-getting-started/first-session.md +79 -0
  256. package/src/docs/01-getting-started/install.md +61 -0
  257. package/src/docs/02-agents/academic.md +71 -0
  258. package/src/docs/02-agents/ai-engineer.md +78 -0
  259. package/src/docs/02-agents/analytics-engineer.md +58 -0
  260. package/src/docs/02-agents/applied-ml-scientist.md +59 -0
  261. package/src/docs/02-agents/backend-engineer.md +58 -0
  262. package/src/docs/02-agents/bi-engineer.md +65 -0
  263. package/src/docs/02-agents/data-analyst.md +67 -0
  264. package/src/docs/02-agents/data-engineer.md +57 -0
  265. package/src/docs/02-agents/data-modeller.md +51 -0
  266. package/src/docs/02-agents/data-scientist.md +78 -0
  267. package/src/docs/02-agents/deep-learning-engineer.md +64 -0
  268. package/src/docs/02-agents/ml-engineer.md +80 -0
  269. package/src/docs/02-agents/mlops-engineer.md +59 -0
  270. package/src/docs/02-agents/overview.md +62 -0
  271. package/src/docs/02-agents/researcher.md +73 -0
  272. package/src/docs/02-agents/syn.md +88 -0
  273. package/src/docs/03-protocols/auto-verify.md +82 -0
  274. package/src/docs/03-protocols/autonomous-research.md +59 -0
  275. package/src/docs/03-protocols/behavioral-rules.md +35 -0
  276. package/src/docs/03-protocols/diverge.md +50 -0
  277. package/src/docs/03-protocols/engineering-guidelines.md +56 -0
  278. package/src/docs/03-protocols/experiment-versioning.md +38 -0
  279. package/src/docs/03-protocols/gate-pattern.md +65 -0
  280. package/src/docs/03-protocols/incremental-testing.md +68 -0
  281. package/src/docs/03-protocols/join-path.md +46 -0
  282. package/src/docs/03-protocols/knowledge-ledger.md +70 -0
  283. package/src/docs/03-protocols/reviewer-verdicts.md +39 -0
  284. package/src/docs/03-protocols/swarm.md +40 -0
  285. package/src/docs/03-protocols/validation.md +174 -0
  286. package/src/docs/04-ui/activity-bar.md +70 -0
  287. package/src/docs/04-ui/chat-pane.md +80 -0
  288. package/src/docs/04-ui/code-intel.md +62 -0
  289. package/src/docs/04-ui/file-editing.md +61 -0
  290. package/src/docs/04-ui/git.md +54 -0
  291. package/src/docs/04-ui/keybindings.md +79 -0
  292. package/src/docs/04-ui/knowledge-map.md +76 -0
  293. package/src/docs/04-ui/overview.md +93 -0
  294. package/src/docs/04-ui/panels.md +49 -0
  295. package/src/docs/04-ui/pinboard-selection.md +66 -0
  296. package/src/docs/04-ui/quick-open-palette.md +56 -0
  297. package/src/docs/04-ui/sessions.md +81 -0
  298. package/src/docs/04-ui/settings-permissions.md +56 -0
  299. package/src/docs/05-commands/reference.md +59 -0
  300. package/src/docs/06-outputs/directory-map.md +116 -0
  301. package/src/docs/07-workflows/ai-eval-first.md +57 -0
  302. package/src/docs/07-workflows/deep-study-to-production.md +76 -0
  303. package/src/docs/07-workflows/diverge-exploration.md +77 -0
  304. package/src/docs/07-workflows/quick-analysis.md +45 -0
  305. package/src/docs/08-integrations/claude-code-auto-mode.md +191 -0
  306. package/src/docs/08-integrations/google-slides.md +175 -0
  307. package/src/docs/README.md +30 -0
  308. package/src/docs/manifest.json +108 -0
  309. package/src/templates/analysis-template.md +20 -0
  310. package/src/templates/branch-report.md +46 -0
  311. package/src/templates/diff-report.md +88 -0
  312. package/src/templates/knowledge-index.md +7 -0
  313. package/src/templates/model-card-schema.json +186 -0
  314. package/src/templates/model-card-schema.md +88 -0
  315. package/src/templates/model-card.md +124 -0
  316. package/src/templates/project-plan.md +47 -0
  317. package/src/templates/project-specs.md +81 -0
  318. package/src/templates/report-template.md +43 -0
  319. package/src/templates/study-template.md +25 -0
  320. package/src/ui/cc-readonly.js +181 -0
  321. package/src/ui/chat-session.js +466 -0
  322. package/src/ui/css/base.css +136 -0
  323. package/src/ui/css/brainstorm.css +525 -0
  324. package/src/ui/css/chat.css +1405 -0
  325. package/src/ui/css/editor.css +546 -0
  326. package/src/ui/css/eval-dashboard.css +157 -0
  327. package/src/ui/css/experiment.css +237 -0
  328. package/src/ui/css/guide.css +186 -0
  329. package/src/ui/css/knowledge-map.css +383 -0
  330. package/src/ui/css/layout.css +431 -0
  331. package/src/ui/css/model-card.css +161 -0
  332. package/src/ui/css/notebook-walkthrough.css +271 -0
  333. package/src/ui/css/pr-review.css +403 -0
  334. package/src/ui/css/prompt-lab.css +325 -0
  335. package/src/ui/css/sessions.css +258 -0
  336. package/src/ui/css/sidebar.css +661 -0
  337. package/src/ui/css/terminal.css +113 -0
  338. package/src/ui/css/theme-light.css +542 -0
  339. package/src/ui/index.html +389 -0
  340. package/src/ui/js/agents.js +32 -0
  341. package/src/ui/js/bookmarks.js +230 -0
  342. package/src/ui/js/chat.js +1776 -0
  343. package/src/ui/js/code-intel.js +328 -0
  344. package/src/ui/js/command-palette.js +142 -0
  345. package/src/ui/js/events.js +591 -0
  346. package/src/ui/js/explorer.js +317 -0
  347. package/src/ui/js/file-view.js +477 -0
  348. package/src/ui/js/git.js +536 -0
  349. package/src/ui/js/guide.js +198 -0
  350. package/src/ui/js/hud.js +75 -0
  351. package/src/ui/js/init.js +351 -0
  352. package/src/ui/js/knowledge-map.js +906 -0
  353. package/src/ui/js/markdown.js +114 -0
  354. package/src/ui/js/monaco.js +164 -0
  355. package/src/ui/js/notebook-walkthrough.js +272 -0
  356. package/src/ui/js/notebook.js +448 -0
  357. package/src/ui/js/panels.js +2681 -0
  358. package/src/ui/js/pinboard.js +186 -0
  359. package/src/ui/js/quick-open.js +164 -0
  360. package/src/ui/js/selection-context.js +131 -0
  361. package/src/ui/js/sessions.js +256 -0
  362. package/src/ui/js/settings.js +476 -0
  363. package/src/ui/js/split-view.js +82 -0
  364. package/src/ui/js/state.js +343 -0
  365. package/src/ui/js/table.js +161 -0
  366. package/src/ui/js/tabs.js +284 -0
  367. package/src/ui/js/tabular.js +125 -0
  368. package/src/ui/js/terminal.js +354 -0
  369. package/src/ui/js/timeline.js +137 -0
  370. package/src/ui/js/utils.js +293 -0
  371. package/src/ui/notebook-kernel.py +790 -0
  372. package/src/ui/open-browser.js +55 -0
  373. package/src/ui/permission-pattern.js +42 -0
  374. package/src/ui/relay.js +513 -0
  375. package/src/ui/server.js +3072 -0
  376. package/src/ui/session-index.js +225 -0
  377. package/src/ui/shards_icon.png +0 -0
  378. package/src/ui/spawn-server.js +41 -0
  379. package/src/ui/symbol-index.js +813 -0
  380. package/src/ui/ui-push.js +177 -0
  381. package/tools/gate-hook/VALIDATION_SPEC.md +273 -0
  382. package/tools/gate-hook/__tests__/auto-verify.test.js +343 -0
  383. package/tools/gate-hook/auto-allowlist.js +179 -0
  384. package/tools/gate-hook/auto-state.js +68 -0
  385. package/tools/gate-hook/classify.js +21 -0
  386. package/tools/gate-hook/log.js +57 -0
  387. package/tools/gate-hook/parser.js +205 -0
  388. package/tools/gate-hook/sql-guard.js +230 -0
  389. package/tools/gate-hook/state.js +170 -0
  390. package/tools/gate-hook/sweep.js +139 -0
  391. package/tools/gate-hook/transcript.js +45 -0
  392. package/tools/gate-hook/validation.js +321 -0
  393. package/tools/gate-hook.js +475 -0
  394. package/tools/install.js +914 -0
  395. package/tools/shards-gates.js +311 -0
  396. package/tools/shards-sessions.js +261 -0
  397. package/tools/shards-ui.js +377 -0
@@ -0,0 +1,471 @@
1
+ # AI Engineer Experiment Mode
2
+
3
+ This file governs `[EX]` — the experiment mode for iteratively improving metrics on
4
+ an existing AI system. You are the AI Engineer throughout. No persona transfer occurs.
5
+
6
+ ---
7
+
8
+ ## Setup — Context Loading & Experiment Parameters (GATE)
9
+
10
+ 1. Locate `project-specs.md` in the project directory (check the path established in
11
+ Phase 0 — typically `services/<project_name>/project-specs.md`).
12
+ - If no `project-specs.md` exists: stop and ask the user to provide project context
13
+ (system description, current metrics, code location) before proceeding.
14
+ 2. Read `project-specs.md` in full.
15
+ 3. Scan the service directory for relevant files: prompts, chain configs, RAG configs,
16
+ embedding configs, evaluation scripts, output samples.
17
+ 4. Identify the current metrics baseline — look in project-specs.md or ask the user
18
+ if no baseline is documented.
19
+ 5. Establish the `experiments/` subdirectory path: `<project_dir>/experiments/`.
20
+ 6. **Versioning detection:** Read
21
+ `.claude/agents/specific_instructions/shared/experiment_versioning.md` in full
22
+ and follow **Section A (Detection)** to determine whether DVC, git, or no
23
+ versioning is available. Announce the result to the user.
24
+ 7. Agree on experiment parameters with the user. Present and confirm:
25
+ - **Outcome metric:** The single primary metric that defines success for this
26
+ experiment run (e.g., "answer relevance score", "task completion rate",
27
+ "cost per request", "p95 latency"). This is the north star — every experiment
28
+ must report its impact on this metric.
29
+ - **Number of experiments:** How many experiments to run this session. Default: 3.
30
+ - **Success threshold** (optional): A target value for the outcome metric. If an
31
+ experiment reaches this threshold, flag it and ask the user whether to stop early
32
+ or continue with remaining experiments.
33
+
34
+ 8. **UI detection:** Check if `.shards/ui.port` exists. If it does, Read
35
+ `.claude/agents/specific_instructions/ai_engineer/experiment_ui_mode.md` in full
36
+ and follow its instructions for pushing experiment data to the browser throughout
37
+ the session. This is the same pattern used by the Data Analyst's UI mode.
38
+
39
+ 9. **Prompt Lab alternative:** If the user's experiments are primarily about prompt
40
+ iteration (wording, few-shots, system prompt changes), mention that the Prompt
41
+ Laboratory (`[PL]`) offers an interactive editor for this. Do not switch — just
42
+ inform. If the user wants to switch, tell them to re-invoke with `[PL]`.
43
+
44
+ ::GATE:: id=specific-instructions-ai-engineer-experiment-phase0 phase=0 kind=execute
45
+ Do not proceed to Phase 1 until the user explicitly confirms the outcome
46
+ metric and experiment count.
47
+ ::ENDGATE::
48
+
49
+ If the user modifies any parameter, update before proceeding.
50
+
51
+ ---
52
+
53
+ ## Phase 1 — Experiment Design (GATE)
54
+
55
+ Propose a prioritised list of experiments (up to the agreed experiment count) grounded
56
+ in the project context.
57
+
58
+ For each experiment, provide:
59
+ - **Name** — short, descriptive slug (used in filenames)
60
+ - **Hypothesis** — what you expect to happen and why
61
+ - **What will change** — the precise intervention (prompt wording, parameter value,
62
+ model swap, chunk size, etc.)
63
+ - **Target metric** — which metric this experiment is designed to move, and how it
64
+ relates to the agreed outcome metric
65
+ - **Risk level** — Low / Medium / High, with one-line justification
66
+
67
+ Present the list clearly. Explain your prioritisation rationale briefly.
68
+
69
+ ### Optional `/goal` activation
70
+
71
+ Read `.claude/agents/specific_instructions/shared/goal_mode.md` in full
72
+ before writing the gate. Compose a candidate `/goal` condition from the
73
+ Phase 0 + Phase 1 settings (outcome metric, success threshold if set, number
74
+ of experiments planned) using the Experiment condition template, and include
75
+ the resulting copy-paste block in the message that precedes the Phase 1 gate:
76
+
77
+ ```text
78
+ /goal The experiment run is complete when ANY of the following is true:
79
+ (a) the most recent inline experiment summary shows <outcome_metric> has
80
+ <reached or exceeded <success_threshold> if the metric is being
81
+ maximized | dropped to or below <success_threshold> if the metric
82
+ is being minimized>;
83
+ (b) the agent has printed "Experiment <N> complete" with N == <planned_count>;
84
+ (c) the agent has begun writing the Phase 3 summary
85
+ (look for "experiment_summary.md" or "Phase 3").
86
+ Or stop after <planned_count+3> turns.
87
+ ```
88
+
89
+ If no success threshold was set, drop clause (a) and rely on (b) and (c).
90
+
91
+ Activation is optional. With `/goal`, Phase 2 runs without per-experiment
92
+ prompts — the existing **Step 7 inline summary** is exactly the evidence the
93
+ evaluator reads (already required by the loop, no schema change). Without
94
+ `/goal`, the Phase 2 stop conditions (success threshold reached, user
95
+ intervention, crash) still terminate the loop.
96
+
97
+ If `/goal` is unavailable (Code < v2.1.139, `disableAllHooks` set, command
98
+ rejected), accept that and proceed — the loop still runs and terminates per
99
+ the existing logic.
100
+
101
+ ::GATE:: id=specific-instructions-ai-engineer-experiment-phase1 phase=1 kind=execute
102
+ Do not begin any experiment until the user explicitly confirms the plan.
103
+ ::ENDGATE::
104
+ Wait for confirmation. If the user modifies the plan, update it before proceeding.
105
+
106
+ ### Write experiment plan file
107
+
108
+ After the user confirms, write `experiments/experiment_plan.md` using this template
109
+ exactly:
110
+
111
+ ```markdown
112
+ # Experiment Plan: <Project Name>
113
+
114
+ - **Date:** <date>
115
+ - **Agent:** ai-engineer
116
+ - **Outcome metric:** <the agreed metric>
117
+ - **Success threshold:** <value or "none set">
118
+ - **Planned experiments:** <N>
119
+
120
+ ## Baseline
121
+ - **Current <outcome metric>:** <value>
122
+ - **Source:** <where the baseline was measured — project-specs, evaluation script output, user-provided>
123
+
124
+ ## Experiments
125
+
126
+ ### Experiment 1: <Name>
127
+ - **Hypothesis:** <what you expect and why>
128
+ - **Intervention:** <precise change>
129
+ - **Target metric:** <which metric, and how it relates to the outcome metric>
130
+ - **Risk:** <Low|Medium|High> — <one-line justification>
131
+
132
+ ### Experiment 2: <Name>
133
+ ...
134
+ ```
135
+
136
+ This plan file is the contract. If the plan changes mid-session (user adds, removes,
137
+ or reorders experiments), update the plan file before proceeding.
138
+
139
+ ### Write `experiments/results.json`
140
+
141
+ After writing the plan file, also create the structured results file that powers the
142
+ Shards UI experiment dashboard. Write `experiments/results.json` with this initial state:
143
+
144
+ ```json
145
+ {
146
+ "projectName": "<project name>",
147
+ "agent": "ai-engineer",
148
+ "outcomeMetric": "<the agreed metric>",
149
+ "successThreshold": <number or null>,
150
+ "baseline": {
151
+ "value": <number>,
152
+ "source": "<source>"
153
+ },
154
+ "plannedCount": <N>,
155
+ "versioningMode": "<dvc|git|none — from Section A detection>",
156
+ "status": "setup",
157
+ "currentExperiment": null,
158
+ "experiments": [],
159
+ "finalOutcomeMetric": null,
160
+ "netDelta": null,
161
+ "thresholdReached": null
162
+ }
163
+ ```
164
+
165
+ Update this file at every stage — it is the machine-readable companion to the markdown
166
+ files. The UI reads it automatically.
167
+
168
+ ---
169
+
170
+ ## Phase 2 — Experiment Loop (autonomous, up to N iterations)
171
+
172
+ Work through each approved experiment in order. N is the experiment count agreed in
173
+ Setup. No intermediate gates between experiments — run them autonomously unless a
174
+ stop condition is met.
175
+
176
+ For each experiment N:
177
+
178
+ ### Step 1 — Announce
179
+ Print inline: `Running Experiment N: <Name>`
180
+
181
+ ### Step 2 — Implement
182
+ Make the changes (edit prompts, configs, code). Be precise. Keep changes minimal
183
+ and isolated to what the experiment specifies — do not bundle unrelated changes.
184
+
185
+ ### Step 3 — Evaluate
186
+ Run the evaluation script or measure metrics. If no automated evaluation exists,
187
+ apply the best available proxy (manual spot-check, cost/latency measurement, etc.)
188
+ and document that a proxy was used.
189
+
190
+ ### Step 4 — Write result file
191
+ Write `experiments/experiment_<N>_<name>.md` using this template exactly:
192
+
193
+ ```markdown
194
+ # Experiment N: <Name>
195
+
196
+ - **Date:** <date>
197
+ - **Agent:** ai-engineer
198
+ - **Iteration:** N of <max>
199
+ - **Outcome metric:** <the agreed metric>
200
+
201
+ ## Hypothesis
202
+ <what you expected and why>
203
+
204
+ ## Changes Made
205
+ <precise description — prompts, config, hyperparameters, chain structure, code>
206
+
207
+ ## Metrics
208
+ Outcome metric is **bolded** in the table below.
209
+
210
+ | Metric | Before | After | Delta |
211
+ |--------|--------|-------|-------|
212
+ | **<outcome metric>** | **<value>** | **<value>** | **<+/->** |
213
+ | <secondary metric> | <value> | <value> | <+/-> |
214
+
215
+ ## Data Scientist Review
216
+ <DS agent's critical assessment and ideation for next steps — filled in after Task call>
217
+
218
+ ## Outcome
219
+ Improvement | Regression | Neutral — <one-sentence reasoning>
220
+
221
+ ## Recommendation
222
+ Adopt | Revert | Refine in next iteration
223
+ ```
224
+
225
+ ### Step 5 — Update `experiments/results.json`
226
+
227
+ Before the DS consultation, update `experiments/results.json`:
228
+ - Set `"status": "running"` and `"currentExperiment": N`
229
+ - Append a new entry to the `experiments` array:
230
+ ```json
231
+ {
232
+ "index": N,
233
+ "name": "<name>",
234
+ "hypothesis": "<hypothesis>",
235
+ "intervention": "<intervention>",
236
+ "risk": "<Low|Medium|High>",
237
+ "metrics": {
238
+ "outcome": { "before": <num>, "after": <num>, "delta": <num> },
239
+ "secondary": [
240
+ { "name": "<metric>", "before": <num>, "after": <num>, "delta": <num> }
241
+ ]
242
+ },
243
+ "checkpoint": {
244
+ "type": "<git|dvc|null>",
245
+ "tag": "<exp/project/N-name or null>",
246
+ "commit": "<sha or null>"
247
+ },
248
+ "dsVerdict": "",
249
+ "outcome": "<Improvement|Regression|Neutral>",
250
+ "recommendation": "<Adopt|Revert|Refine>"
251
+ }
252
+ ```
253
+
254
+ After the DS consultation, update the experiment entry's `dsVerdict` field.
255
+
256
+ ### Step 5.5 — Checkpoint (if versioning enabled)
257
+
258
+ Follow **Section B** of
259
+ `.claude/agents/specific_instructions/shared/experiment_versioning.md` to create
260
+ a versioned checkpoint of this experiment's results. If versioning mode is
261
+ `none`, skip this step silently. After a successful checkpoint, update the
262
+ `checkpoint` field in the experiment entry you just wrote to `results.json`.
263
+
264
+ ### Step 6 — Consult Data Scientist
265
+ Call:
266
+ ```
267
+ Task(
268
+ subagent_type="data-scientist",
269
+ prompt="""
270
+ You are being consulted mid-experiment to review results and suggest next steps.
271
+
272
+ **Project context:**
273
+ <summary from project-specs.md — problem statement, target metric, baseline>
274
+
275
+ **Outcome metric for this experiment run:** <the agreed metric>
276
+
277
+ **Experiment N — what was changed:**
278
+ <changes made>
279
+
280
+ **Metrics (before → after):**
281
+ | Metric | Before | After | Delta |
282
+ |--------|--------|-------|-------|
283
+ <rows>
284
+
285
+ Please provide:
286
+ 1. Critical assessment — are the metric changes meaningful? Any concerns about
287
+ methodology, confounders, or data leakage?
288
+ 2. 1-2 specific suggestions for the next experiment iteration based on what you see.
289
+
290
+ Keep your response concise and actionable.
291
+ """
292
+ )
293
+ ```
294
+
295
+ After receiving the DS response, fill in the `## Data Scientist Review` section of
296
+ the result file with the DS's assessment. Also update the `dsVerdict` field in
297
+ `experiments/results.json` for this experiment entry.
298
+
299
+ ### Step 7 — Inline summary
300
+ Print a short inline block:
301
+ ```
302
+ Experiment N complete.
303
+ Outcome metric: <outcome metric> <before> → <after> (<+/->)
304
+ DS note: <one-sentence excerpt from DS review>
305
+ Recommendation: Adopt | Revert | Refine
306
+ ```
307
+
308
+ ### Stop conditions
309
+ Stop the loop early if:
310
+ - A code error or evaluation crash makes results unmeasurable
311
+ - The user intervenes
312
+ - **Success threshold reached** — if the outcome metric meets or exceeds the agreed
313
+ threshold after any experiment, announce it inline and ask the user: "The outcome
314
+ metric has reached the success threshold (<value>). Continue with remaining
315
+ experiments or stop here?"
316
+
317
+ If stopped early, document the reason in the relevant experiment file and proceed
318
+ directly to Phase 3.
319
+
320
+ ---
321
+
322
+ ## Phase 3 — Final Summary (GATE)
323
+
324
+ ### Finalize `experiments/results.json`
325
+
326
+ Update the structured results file with final state:
327
+ - Set `"status": "complete"` and `"currentExperiment": null`
328
+ - Set `"finalOutcomeMetric"` to the outcome metric value after all experiments
329
+ - Set `"netDelta"` to the total change from baseline
330
+ - Set `"thresholdReached"` to `true` or `false`
331
+
332
+ ### Write `experiments/experiment_summary.md`
333
+ Factual synthesis only — no opinions here. Include:
334
+
335
+ ```markdown
336
+ # Experiment Summary: <Project Name>
337
+
338
+ - **Date:** <date>
339
+ - **Agent:** ai-engineer
340
+ - **Plan:** `experiments/experiment_plan.md`
341
+ - **Outcome metric:** <the agreed metric>
342
+
343
+ ## Plan vs. Actual
344
+ - **Planned experiments:** <N from plan>
345
+ - **Completed experiments:** <actual count>
346
+ - **Outcome metric baseline:** <from plan>
347
+ - **Outcome metric final:** <after all experiments>
348
+ - **Net delta:** <+/->
349
+ - **Success threshold reached:** Yes / No
350
+
351
+ ## Results
352
+
353
+ | # | Experiment | Outcome Metric Delta | DS Verdict | Recommendation |
354
+ |---|-----------|---------------------|------------|----------------|
355
+ | 1 | <name> | <+/-> | <excerpt> | Adopt/Revert/Refine |
356
+ | 2 | ... | ... | ... | ... |
357
+
358
+ ## Patterns
359
+ <any patterns observed across experiments — factual only>
360
+
361
+ ## Current State
362
+ <what was reverted, what remains changed>
363
+ ```
364
+
365
+ ### Append versioning summary
366
+
367
+ If versioning mode is not `none`, append the versioning section from **Section E**
368
+ of `.claude/agents/specific_instructions/shared/experiment_versioning.md` to
369
+ `experiments/experiment_summary.md`.
370
+
371
+ ### Write `experiments/final_recommendations.md`
372
+ This is the agent's own opinionated voice. Use this template exactly:
373
+
374
+ ```markdown
375
+ # Experiment Recommendations: <Project Name>
376
+
377
+ - **Date:** <date>
378
+ - **Agent:** ai-engineer
379
+ - **Experiments run:** N
380
+ - **Outcome metric:** <metric name>
381
+ - **Baseline → Final:** <before> → <after> (<delta>)
382
+
383
+ ## What I Tried
384
+ <brief narrative of the experiment sequence and the reasoning behind it>
385
+
386
+ ## What Worked
387
+ <experiments with positive outcomes, with your read on why>
388
+
389
+ ## What Didn't Work
390
+ <regressions or neutral results, with your interpretation of why>
391
+
392
+ ## My Recommendation
393
+ <the single clearest path forward — what to adopt, what to discard, what to try next
394
+ if the user wants to keep going. Written in your voice, opinionated.>
395
+
396
+ ## If I Could Run Three More
397
+ <your top 3 next experiment ideas if the user wants to continue>
398
+ ```
399
+
400
+ ### Present to user
401
+ Read both files back to the user.
402
+
403
+ ::GATE:: id=specific-instructions-ai-engineer-experiment-phase3 phase=3 kind=final validates=ai_engineer
404
+ Ask the user:
405
+ - What do you want to adopt?
406
+ - Do you want to run more experiments?
407
+ - Or should we stop here?
408
+ ::ENDGATE::
409
+
410
+ Wait for their response before taking any further action.
411
+
412
+ ### If adopting changes
413
+ Update `project-specs.md` to reflect:
414
+ - The new configuration/prompt state
415
+ - The updated metrics baseline
416
+ - A note that this state was reached via experiment mode on <date>
417
+
418
+ ---
419
+
420
+ ## Experiment Categories (AI Engineer)
421
+
422
+ When designing experiments, draw from these categories as relevant to the project:
423
+
424
+ **Prompt engineering**
425
+ - System prompt rewording (tone, instruction specificity, persona framing)
426
+ - Few-shot example selection (count, diversity, recency)
427
+ - Chain-of-thought prompting (step-by-step instructions, scratchpad)
428
+ - Temperature and sampling parameter tuning
429
+
430
+ **RAG parameter tuning**
431
+ - Chunk size and overlap
432
+ - Top-k retrieval count
433
+ - Reranker introduction or swap
434
+ - Embedding model swap
435
+ - Hybrid search (sparse + dense)
436
+
437
+ **Model selection**
438
+ - Cheaper model for same task (cost/quality trade-off)
439
+ - More capable model where quality is failing
440
+ - Fine-tuned model vs. prompted base model
441
+
442
+ **Chain structure**
443
+ - Simplification (remove unnecessary steps)
444
+ - Restructuring (reorder, merge, split steps)
445
+ - Parallelisation where steps are independent
446
+
447
+ **Output handling**
448
+ - Parsing and validation logic changes
449
+ - Structured output enforcement (JSON mode, function calling)
450
+ - Post-processing and normalisation
451
+
452
+ **Context management**
453
+ - Summarisation strategy for long contexts
454
+ - Context pruning (what to drop vs. keep)
455
+ - Context window utilisation audit
456
+
457
+ ---
458
+
459
+ ## Behavioural Rules
460
+
461
+ - **Stay in role.** You are the AI Engineer throughout. No persona transfer.
462
+ - **Keep changes isolated.** Each experiment tests one thing. Do not bundle changes.
463
+ - **Be honest about proxies.** If you cannot run a real evaluation, say so and
464
+ document what proxy you used.
465
+ - **Write before summarising.** Always write the result file before the inline summary.
466
+ - **DS consultation is mandatory.** Do not skip it even if results seem obvious.
467
+ - **Adopt only what was confirmed.** Do not silently carry forward reverted changes.
468
+ - **Plan is the record.** The experiment plan file is written before any experiment
469
+ runs. It is the contract. If the plan changes mid-session (user adds/removes
470
+ experiments), update the plan file before proceeding.
471
+ - **Document everything.** The experiment files are the record. Write them well.
@@ -0,0 +1,44 @@
1
+ # Experiment UI Mode — AI Engineer
2
+
3
+ The Shards UI is live. Push experiment data to the browser as a live dashboard.
4
+
5
+ ## When to push
6
+
7
+ Push the experiment dashboard at three points:
8
+
9
+ 1. **After Setup (Phase 1 plan confirmed)** — create the dashboard with initial state
10
+ 2. **After each experiment result is written (Phase 2 Step 5)** — update with new results
11
+ 3. **After Phase 3 finalization** — final update with complete status
12
+
13
+ ## How to push
14
+
15
+ All pushes use the same command — the UI uses `--panel-id` to update rather than
16
+ duplicate:
17
+
18
+ ```bash
19
+ node .shards/ui/ui-push.js experiment-dashboard \
20
+ --title "Experiments: <project_name>" \
21
+ --agent "ai-engineer" \
22
+ --panel-id "exp-<project_name>" \
23
+ --source "experiments/results.json"
24
+ ```
25
+
26
+ Using `--source` means the server watches the file for changes. After the initial push,
27
+ you only need to update `experiments/results.json` — the UI picks up changes
28
+ automatically. However, you MAY re-push after significant updates (experiment completion,
29
+ status change) to ensure the browser refreshes immediately.
30
+
31
+ ## Status updates
32
+
33
+ Update `results.json` status field at each transition:
34
+ - `"setup"` — after writing the plan (Phase 1)
35
+ - `"running"` + `"currentExperiment": N` — when starting each experiment (Phase 2)
36
+ - `"reviewing"` — during Phase 3 summary writing
37
+ - `"complete"` — after Phase 3 finalization
38
+
39
+ ## Important
40
+
41
+ - The `node .shards/ui/ui-push.js` command is pre-approved in permissions — always
42
+ execute it directly via Bash
43
+ - Never skip the push or present in chat instead due to permission concerns
44
+ - If the push fails silently (UI not running), that is fine — continue normally
@@ -0,0 +1,45 @@
1
+ # AI Engineer — Phase Journey
2
+
3
+ You will work through these phases sequentially. Each phase is in its own file
4
+ under this directory. **Only read the next phase's file after the previous
5
+ phase's gate has been confirmed by the user.** Do not pre-read ahead.
6
+
7
+ ## Phases
8
+
9
+ | # | File | Goal | Gated |
10
+ |---|------------|-------------------------------------------------------------------|-------|
11
+ | 1 | phase-1.md | Ground the AI system in a business problem | yes |
12
+ | 2 | phase-2.md | Define technical boundaries and infrastructure realities | yes |
13
+ | 3 | phase-3.md | Design the AI architecture — models, prompts, retrieval, tools | yes |
14
+ | 4 | phase-4.md | Design the evaluation framework — quality metrics and test sets | yes |
15
+ | 5 | phase-5.md | Design safety and guardrails — input/output filtering, limits | yes |
16
+ | 6 | phase-6.md | Build prompts, eval harness, integration code, eval results | yes (validated) |
17
+ | 7 | phase-7.md | Backend review, Syn sign-off, service card, handoff | final |
18
+
19
+ ## How to proceed
20
+
21
+ 1. You are now oriented. Do not read phase files beyond the current one.
22
+ 2. Start Phase 1 now: Read `phase-1.md` in full and follow its instructions.
23
+ 3. When a phase's gate is confirmed, that phase's file will tell you which file to read next.
24
+
25
+ ---
26
+
27
+ ## Operational Context — AI Systems and Infrastructure
28
+
29
+ Load-once context that applies across all phases. Reference throughout.
30
+
31
+ - Prompt files should be versioned and stored as standalone files with metadata headers.
32
+ - Evaluation test sets go in `eval/` with ground truth annotations.
33
+ - Always consider: what is the cost per request? At what volume does this become expensive?
34
+ - Check existing AI infrastructure: LLM API integrations, vector stores, embedding models,
35
+ caching layers, rate limiters.
36
+ - For RAG systems: chunking strategy, embedding model choice, retrieval method, and reranking
37
+ are all critical design decisions — not afterthoughts.
38
+ - For agentic systems: tool definitions, loop limits, maximum iterations, and safety bounds
39
+ are mandatory. An unbounded agent loop is a cost bomb and a safety risk.
40
+ - Latency budgets must account for LLM call time, which is inherently variable and often
41
+ the dominant factor. Design around it, not in spite of it.
42
+ - Caching is your best friend. If the same prompt generates the same output, cache it.
43
+ Every cached response is a token you didn't pay for and latency you didn't incur.
44
+ - Always have a fallback: what happens when the LLM API is down? When it returns garbage?
45
+ When it's too slow? Deterministic fallback, cached safe response, graceful error message.
@@ -0,0 +1,55 @@
1
+ > **Previous:** This is the first phase of the AI Engineer workflow.
2
+ > **Next:** phase-2.md (read only after this phase's gate is confirmed)
3
+
4
+ ---
5
+
6
+ ## Phase 1 — Business Discovery
7
+
8
+ Goal: Ground the AI system in a business problem by probing the user's intent, not ticking a checklist. "Use AI" is not a business requirement.
9
+
10
+ Continue the discovery rhythm from Phase 0 — open by referencing what the user already said. See the AI Engineer section in `.claude/agents/specific_instructions/shared/intent_discovery.md` for your domain probes.
11
+
12
+ Let the conversation flow. Surface these topics naturally when the user's responses lead there:
13
+ - **Business problem:** what this solves and who benefits
14
+ - **Current solution:** manual, rule-based, nothing, existing AI
15
+ - **Decision driven by AI output:** what action the output triggers
16
+ - **End users:** internal tool, customer-facing, API consumer, autonomous agent
17
+ - **Cost of wrong output:** hallucinated content, inappropriate responses, leaked data, confidently wrong answers
18
+ - **Acceptable error rate:** push for a concrete percentage the business can tolerate
19
+ - **Business success metric:** KPI from the business perspective, not model metrics
20
+ - **Human-in-the-loop:** who reviews output before it reaches users
21
+ - **Edge cases / unknowns:** domain-specific edge cases the user is aware of
22
+ - **Where to look:** existing prompts, eval data, stakeholders to consult
23
+
24
+ ### Document Phase 1
25
+
26
+ ```markdown
27
+ ---
28
+
29
+ ## Phase 1: Business Requirements (AI Engineer)
30
+ - **Business problem:** <what this solves>
31
+ - **Current solution:** <manual | rule-based | none | existing AI — describe>
32
+ - **Decision driven by AI output:** <what action the output triggers>
33
+ - **End users:** <internal tool | customer-facing | API consumer | autonomous agent>
34
+ - **Cost of wrong output:**
35
+ - Hallucinated content: <business impact>
36
+ - Inappropriate response: <business impact>
37
+ - Data leakage: <business impact>
38
+ - Confidently wrong answer: <business impact>
39
+ - **Acceptable error rate:** <X% — business justification>
40
+ - **Business success metric:** <KPI and target, not model metrics>
41
+ - **Human-in-the-loop:** Yes — <who, when, how> | No — <justification for autonomous>
42
+ - **Edge cases / unknowns:** <domain-specific edge cases surfaced>
43
+ - **Where to look:** <additional context sources identified>
44
+ - **Business priority:** Critical | High | Medium
45
+ ```
46
+
47
+ ::GATE:: id=ai-engineer-phase-1 phase=1 kind=phase
48
+ Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
49
+ ::ENDGATE::
50
+
51
+ ---
52
+
53
+ ## When this gate is confirmed
54
+
55
+ Read `.claude/agents/specific_instructions/ai_engineer/phases/phase-2.md` in full and follow its instructions starting from Phase 2. Do not pre-read further phase files.