@proflandrigan/shards 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +475 -0
- package/package.json +37 -0
- package/src/agents/academic.md +276 -0
- package/src/agents/ai-engineer.md +377 -0
- package/src/agents/analytics-engineer.md +364 -0
- package/src/agents/applied-ml-scientist.md +410 -0
- package/src/agents/backend-engineer.md +255 -0
- package/src/agents/bi-engineer.md +333 -0
- package/src/agents/data-analyst.md +343 -0
- package/src/agents/data-engineer.md +260 -0
- package/src/agents/data-modeller.md +386 -0
- package/src/agents/data-scientist.md +366 -0
- package/src/agents/deep-learning-engineer.md +389 -0
- package/src/agents/ml-engineer.md +424 -0
- package/src/agents/mlops-engineer.md +339 -0
- package/src/agents/researcher.md +187 -0
- package/src/agents/specific_instructions/academic/critical_review.md +263 -0
- package/src/agents/specific_instructions/academic/report.md +113 -0
- package/src/agents/specific_instructions/ai_engineer/advise.md +162 -0
- package/src/agents/specific_instructions/ai_engineer/bi_engineer_handoff.md +86 -0
- package/src/agents/specific_instructions/ai_engineer/experiment.md +471 -0
- package/src/agents/specific_instructions/ai_engineer/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/ai_engineer/phases/index.md +45 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-1.md +55 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-2.md +86 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-3.md +96 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-4.md +138 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-5.md +157 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-6.md +196 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-7.md +313 -0
- package/src/agents/specific_instructions/ai_engineer/phases.md +1011 -0
- package/src/agents/specific_instructions/ai_engineer/prompt_lab.md +161 -0
- package/src/agents/specific_instructions/ai_engineer/prompt_lab_ui_mode.md +28 -0
- package/src/agents/specific_instructions/ai_engineer/research.md +393 -0
- package/src/agents/specific_instructions/ai_engineer/research_ui_mode.md +66 -0
- package/src/agents/specific_instructions/ai_engineer/review.md +159 -0
- package/src/agents/specific_instructions/ai_engineer/validation_checklist.md +182 -0
- package/src/agents/specific_instructions/analytics_engineer/advise.md +155 -0
- package/src/agents/specific_instructions/analytics_engineer/bi_engineer_handoff.md +91 -0
- package/src/agents/specific_instructions/analytics_engineer/data_analyst_handoff.md +84 -0
- package/src/agents/specific_instructions/analytics_engineer/deep_phases.md +818 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/index.md +24 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-1.md +77 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-2.md +106 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-3.md +93 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-4.md +79 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-5.md +61 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-6.md +45 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-7.md +235 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-8.md +221 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-1.md +47 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-2.md +78 -0
- package/src/agents/specific_instructions/analytics_engineer/quick_phases.md +112 -0
- package/src/agents/specific_instructions/analytics_engineer/review.md +167 -0
- package/src/agents/specific_instructions/analytics_engineer/service_mode.md +369 -0
- package/src/agents/specific_instructions/analytics_engineer/ui_mode.md +45 -0
- package/src/agents/specific_instructions/analytics_engineer/update.md +162 -0
- package/src/agents/specific_instructions/analytics_engineer/validation_checklist.md +121 -0
- package/src/agents/specific_instructions/applied_ml_scientist/advise.md +143 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/index.md +21 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-1.md +51 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-2.md +66 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-3.md +113 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-4.md +104 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-5.md +156 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases.md +428 -0
- package/src/agents/specific_instructions/applied_ml_scientist/research.md +379 -0
- package/src/agents/specific_instructions/applied_ml_scientist/review.md +142 -0
- package/src/agents/specific_instructions/applied_ml_scientist/validation_checklist.md +136 -0
- package/src/agents/specific_instructions/backend_engineer/clean.md +149 -0
- package/src/agents/specific_instructions/backend_engineer/review.md +91 -0
- package/src/agents/specific_instructions/backend_engineer/review_checklist.md +54 -0
- package/src/agents/specific_instructions/backend_engineer/service_mode.md +67 -0
- package/src/agents/specific_instructions/bi_engineer/advise.md +137 -0
- package/src/agents/specific_instructions/bi_engineer/data_analyst_handoff.md +77 -0
- package/src/agents/specific_instructions/bi_engineer/incoming_handoff.md +45 -0
- package/src/agents/specific_instructions/bi_engineer/phases/index.md +20 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-1.md +164 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-2.md +92 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-3.md +121 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-4.md +106 -0
- package/src/agents/specific_instructions/bi_engineer/phases.md +451 -0
- package/src/agents/specific_instructions/bi_engineer/review.md +166 -0
- package/src/agents/specific_instructions/bi_engineer/update.md +147 -0
- package/src/agents/specific_instructions/bi_engineer/validation_checklist.md +124 -0
- package/src/agents/specific_instructions/data_analyst/advise.md +138 -0
- package/src/agents/specific_instructions/data_analyst/explain.md +221 -0
- package/src/agents/specific_instructions/data_analyst/incoming_handoff.md +40 -0
- package/src/agents/specific_instructions/data_analyst/phases/index.md +20 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-1.md +159 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-2.md +112 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-3.md +265 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-4.md +100 -0
- package/src/agents/specific_instructions/data_analyst/phases.md +501 -0
- package/src/agents/specific_instructions/data_analyst/review.md +138 -0
- package/src/agents/specific_instructions/data_analyst/ui_mode.md +26 -0
- package/src/agents/specific_instructions/data_analyst/update.md +144 -0
- package/src/agents/specific_instructions/data_analyst/validation_checklist.md +95 -0
- package/src/agents/specific_instructions/data_engineer/advise.md +137 -0
- package/src/agents/specific_instructions/data_engineer/phases.md +466 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/index.md +23 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-1.md +49 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-2.md +93 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-3.md +55 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-4.md +48 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-5.md +40 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-6.md +102 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-7.md +87 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/phase-1.md +45 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/phase-2.md +54 -0
- package/src/agents/specific_instructions/data_engineer/review.md +135 -0
- package/src/agents/specific_instructions/data_engineer/validation_checklist.md +136 -0
- package/src/agents/specific_instructions/data_modeller/advise.md +137 -0
- package/src/agents/specific_instructions/data_modeller/phases.md +581 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/index.md +23 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-1.md +52 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-2.md +113 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-3.md +47 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-4.md +51 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-5.md +45 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-6.md +105 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-7.md +136 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/phase-1.md +47 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/phase-2.md +65 -0
- package/src/agents/specific_instructions/data_modeller/review.md +141 -0
- package/src/agents/specific_instructions/data_modeller/service_mode.md +218 -0
- package/src/agents/specific_instructions/data_modeller/validation_checklist.md +125 -0
- package/src/agents/specific_instructions/data_scientist/advise.md +158 -0
- package/src/agents/specific_instructions/data_scientist/bi_engineer_handoff.md +63 -0
- package/src/agents/specific_instructions/data_scientist/experiment.md +482 -0
- package/src/agents/specific_instructions/data_scientist/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/data_scientist/explain.md +247 -0
- package/src/agents/specific_instructions/data_scientist/greenfield_data.md +35 -0
- package/src/agents/specific_instructions/data_scientist/ml_engineer_handoff.md +52 -0
- package/src/agents/specific_instructions/data_scientist/notebook_walkthrough.md +76 -0
- package/src/agents/specific_instructions/data_scientist/phases/index.md +24 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-1.md +45 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-2.md +67 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-3.md +89 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-4.md +143 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-5.md +71 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-6.md +239 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-7.md +207 -0
- package/src/agents/specific_instructions/data_scientist/phases.md +651 -0
- package/src/agents/specific_instructions/data_scientist/research.md +345 -0
- package/src/agents/specific_instructions/data_scientist/research_ui_mode.md +52 -0
- package/src/agents/specific_instructions/data_scientist/review.md +136 -0
- package/src/agents/specific_instructions/data_scientist/service_mode.md +247 -0
- package/src/agents/specific_instructions/data_scientist/validation_checklist.md +183 -0
- package/src/agents/specific_instructions/deep_learning_engineer/advise.md +145 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/index.md +21 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-1.md +74 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-2.md +98 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-3.md +76 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-4.md +128 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-5.md +292 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases.md +567 -0
- package/src/agents/specific_instructions/deep_learning_engineer/research.md +389 -0
- package/src/agents/specific_instructions/deep_learning_engineer/review.md +155 -0
- package/src/agents/specific_instructions/deep_learning_engineer/validation_checklist.md +147 -0
- package/src/agents/specific_instructions/ml_engineer/advise.md +174 -0
- package/src/agents/specific_instructions/ml_engineer/bi_engineer_handoff.md +71 -0
- package/src/agents/specific_instructions/ml_engineer/experiment.md +474 -0
- package/src/agents/specific_instructions/ml_engineer/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/ml_engineer/notebook_walkthrough.md +75 -0
- package/src/agents/specific_instructions/ml_engineer/phases/index.md +25 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-1.md +49 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-2.md +75 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-3.md +124 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-4.md +279 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-5.md +160 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-6-5.md +170 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-6.md +295 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-7.md +337 -0
- package/src/agents/specific_instructions/ml_engineer/phases.md +1068 -0
- package/src/agents/specific_instructions/ml_engineer/research.md +437 -0
- package/src/agents/specific_instructions/ml_engineer/research_ui_mode.md +71 -0
- package/src/agents/specific_instructions/ml_engineer/review.md +187 -0
- package/src/agents/specific_instructions/ml_engineer/service_mode.md +273 -0
- package/src/agents/specific_instructions/ml_engineer/validation_checklist.md +185 -0
- package/src/agents/specific_instructions/mlops_engineer/advise.md +139 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/index.md +23 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-1.md +52 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-2.md +86 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-3.md +105 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-4.md +128 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-5.md +106 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-6.md +128 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-7.md +144 -0
- package/src/agents/specific_instructions/mlops_engineer/phases.md +671 -0
- package/src/agents/specific_instructions/mlops_engineer/review.md +164 -0
- package/src/agents/specific_instructions/mlops_engineer/service_mode.md +81 -0
- package/src/agents/specific_instructions/mlops_engineer/validation_checklist.md +151 -0
- package/src/agents/specific_instructions/researcher/critical_review.md +292 -0
- package/src/agents/specific_instructions/researcher/review_checklist.md +67 -0
- package/src/agents/specific_instructions/researcher/service_mode.md +224 -0
- package/src/agents/specific_instructions/shared/auto_verify_mode.md +141 -0
- package/src/agents/specific_instructions/shared/autonomous_research.md +1289 -0
- package/src/agents/specific_instructions/shared/behavioral_rules.md +36 -0
- package/src/agents/specific_instructions/shared/diverge_protocol.md +387 -0
- package/src/agents/specific_instructions/shared/engineering_guidelines.md +136 -0
- package/src/agents/specific_instructions/shared/experiment_versioning.md +184 -0
- package/src/agents/specific_instructions/shared/goal_mode.md +187 -0
- package/src/agents/specific_instructions/shared/incremental_testing.md +139 -0
- package/src/agents/specific_instructions/shared/intent_discovery.md +223 -0
- package/src/agents/specific_instructions/shared/join_path_protocol.md +168 -0
- package/src/agents/specific_instructions/shared/knowledge_checkpoint.md +83 -0
- package/src/agents/specific_instructions/shared/knowledge_harvest.md +220 -0
- package/src/agents/specific_instructions/shared/knowledge_retrieval.md +100 -0
- package/src/agents/specific_instructions/shared/notebook_walkthrough_protocol.md +367 -0
- package/src/agents/specific_instructions/shared/reviewer_verdict_protocol.md +74 -0
- package/src/agents/specific_instructions/shared/swarm_protocol.md +97 -0
- package/src/agents/specific_instructions/shared/validation_protocol.md +139 -0
- package/src/agents/specific_instructions/syn/arbiter.md +140 -0
- package/src/agents/specific_instructions/syn/brainstorm.md +550 -0
- package/src/agents/specific_instructions/syn/code_review.md +232 -0
- package/src/agents/specific_instructions/syn/diff.md +239 -0
- package/src/agents/specific_instructions/syn/final_review.md +65 -0
- package/src/agents/specific_instructions/syn/fixer.md +240 -0
- package/src/agents/specific_instructions/syn/free_form.md +130 -0
- package/src/agents/specific_instructions/syn/knowledge.md +468 -0
- package/src/agents/specific_instructions/syn/notebook_walkthrough.md +78 -0
- package/src/agents/specific_instructions/syn/panel_review.md +634 -0
- package/src/agents/specific_instructions/syn/pm.md +453 -0
- package/src/agents/specific_instructions/syn/pr_review.md +255 -0
- package/src/agents/specific_instructions/syn/slides.md +417 -0
- package/src/agents/syn.md +729 -0
- package/src/commands/academic.md +41 -0
- package/src/commands/ai-engineer.md +45 -0
- package/src/commands/analytics-engineer.md +48 -0
- package/src/commands/applied-ml-scientist.md +45 -0
- package/src/commands/backend-engineer.md +35 -0
- package/src/commands/bi-engineer.md +40 -0
- package/src/commands/brainstorm.md +24 -0
- package/src/commands/data-analyst.md +38 -0
- package/src/commands/data-engineer.md +37 -0
- package/src/commands/data-modeller.md +38 -0
- package/src/commands/data-scientist.md +38 -0
- package/src/commands/deep-learning-engineer.md +47 -0
- package/src/commands/end.md +49 -0
- package/src/commands/knowledge.md +24 -0
- package/src/commands/ml-engineer.md +42 -0
- package/src/commands/mlops-engineer.md +47 -0
- package/src/commands/notebook-walkthrough.md +58 -0
- package/src/commands/researcher.md +40 -0
- package/src/commands/resume.md +57 -0
- package/src/commands/review-pr.md +26 -0
- package/src/commands/shards-guide.md +41 -0
- package/src/commands/shards-ui.md +32 -0
- package/src/commands/shards.md +41 -0
- package/src/docs/01-getting-started/concepts.md +109 -0
- package/src/docs/01-getting-started/first-session.md +79 -0
- package/src/docs/01-getting-started/install.md +61 -0
- package/src/docs/02-agents/academic.md +71 -0
- package/src/docs/02-agents/ai-engineer.md +78 -0
- package/src/docs/02-agents/analytics-engineer.md +58 -0
- package/src/docs/02-agents/applied-ml-scientist.md +59 -0
- package/src/docs/02-agents/backend-engineer.md +58 -0
- package/src/docs/02-agents/bi-engineer.md +65 -0
- package/src/docs/02-agents/data-analyst.md +67 -0
- package/src/docs/02-agents/data-engineer.md +57 -0
- package/src/docs/02-agents/data-modeller.md +51 -0
- package/src/docs/02-agents/data-scientist.md +78 -0
- package/src/docs/02-agents/deep-learning-engineer.md +64 -0
- package/src/docs/02-agents/ml-engineer.md +80 -0
- package/src/docs/02-agents/mlops-engineer.md +59 -0
- package/src/docs/02-agents/overview.md +62 -0
- package/src/docs/02-agents/researcher.md +73 -0
- package/src/docs/02-agents/syn.md +88 -0
- package/src/docs/03-protocols/auto-verify.md +82 -0
- package/src/docs/03-protocols/autonomous-research.md +59 -0
- package/src/docs/03-protocols/behavioral-rules.md +35 -0
- package/src/docs/03-protocols/diverge.md +50 -0
- package/src/docs/03-protocols/engineering-guidelines.md +56 -0
- package/src/docs/03-protocols/experiment-versioning.md +38 -0
- package/src/docs/03-protocols/gate-pattern.md +65 -0
- package/src/docs/03-protocols/incremental-testing.md +68 -0
- package/src/docs/03-protocols/join-path.md +46 -0
- package/src/docs/03-protocols/knowledge-ledger.md +70 -0
- package/src/docs/03-protocols/reviewer-verdicts.md +39 -0
- package/src/docs/03-protocols/swarm.md +40 -0
- package/src/docs/03-protocols/validation.md +174 -0
- package/src/docs/04-ui/activity-bar.md +70 -0
- package/src/docs/04-ui/chat-pane.md +80 -0
- package/src/docs/04-ui/code-intel.md +62 -0
- package/src/docs/04-ui/file-editing.md +61 -0
- package/src/docs/04-ui/git.md +54 -0
- package/src/docs/04-ui/keybindings.md +79 -0
- package/src/docs/04-ui/knowledge-map.md +76 -0
- package/src/docs/04-ui/overview.md +93 -0
- package/src/docs/04-ui/panels.md +49 -0
- package/src/docs/04-ui/pinboard-selection.md +66 -0
- package/src/docs/04-ui/quick-open-palette.md +56 -0
- package/src/docs/04-ui/sessions.md +81 -0
- package/src/docs/04-ui/settings-permissions.md +56 -0
- package/src/docs/05-commands/reference.md +59 -0
- package/src/docs/06-outputs/directory-map.md +116 -0
- package/src/docs/07-workflows/ai-eval-first.md +57 -0
- package/src/docs/07-workflows/deep-study-to-production.md +76 -0
- package/src/docs/07-workflows/diverge-exploration.md +77 -0
- package/src/docs/07-workflows/quick-analysis.md +45 -0
- package/src/docs/08-integrations/claude-code-auto-mode.md +191 -0
- package/src/docs/08-integrations/google-slides.md +175 -0
- package/src/docs/README.md +30 -0
- package/src/docs/manifest.json +108 -0
- package/src/templates/analysis-template.md +20 -0
- package/src/templates/branch-report.md +46 -0
- package/src/templates/diff-report.md +88 -0
- package/src/templates/knowledge-index.md +7 -0
- package/src/templates/model-card-schema.json +186 -0
- package/src/templates/model-card-schema.md +88 -0
- package/src/templates/model-card.md +124 -0
- package/src/templates/project-plan.md +47 -0
- package/src/templates/project-specs.md +81 -0
- package/src/templates/report-template.md +43 -0
- package/src/templates/study-template.md +25 -0
- package/src/ui/cc-readonly.js +181 -0
- package/src/ui/chat-session.js +466 -0
- package/src/ui/css/base.css +136 -0
- package/src/ui/css/brainstorm.css +525 -0
- package/src/ui/css/chat.css +1405 -0
- package/src/ui/css/editor.css +546 -0
- package/src/ui/css/eval-dashboard.css +157 -0
- package/src/ui/css/experiment.css +237 -0
- package/src/ui/css/guide.css +186 -0
- package/src/ui/css/knowledge-map.css +383 -0
- package/src/ui/css/layout.css +431 -0
- package/src/ui/css/model-card.css +161 -0
- package/src/ui/css/notebook-walkthrough.css +271 -0
- package/src/ui/css/pr-review.css +403 -0
- package/src/ui/css/prompt-lab.css +325 -0
- package/src/ui/css/sessions.css +258 -0
- package/src/ui/css/sidebar.css +661 -0
- package/src/ui/css/terminal.css +113 -0
- package/src/ui/css/theme-light.css +542 -0
- package/src/ui/index.html +389 -0
- package/src/ui/js/agents.js +32 -0
- package/src/ui/js/bookmarks.js +230 -0
- package/src/ui/js/chat.js +1776 -0
- package/src/ui/js/code-intel.js +328 -0
- package/src/ui/js/command-palette.js +142 -0
- package/src/ui/js/events.js +591 -0
- package/src/ui/js/explorer.js +317 -0
- package/src/ui/js/file-view.js +477 -0
- package/src/ui/js/git.js +536 -0
- package/src/ui/js/guide.js +198 -0
- package/src/ui/js/hud.js +75 -0
- package/src/ui/js/init.js +351 -0
- package/src/ui/js/knowledge-map.js +906 -0
- package/src/ui/js/markdown.js +114 -0
- package/src/ui/js/monaco.js +164 -0
- package/src/ui/js/notebook-walkthrough.js +272 -0
- package/src/ui/js/notebook.js +448 -0
- package/src/ui/js/panels.js +2681 -0
- package/src/ui/js/pinboard.js +186 -0
- package/src/ui/js/quick-open.js +164 -0
- package/src/ui/js/selection-context.js +131 -0
- package/src/ui/js/sessions.js +256 -0
- package/src/ui/js/settings.js +476 -0
- package/src/ui/js/split-view.js +82 -0
- package/src/ui/js/state.js +343 -0
- package/src/ui/js/table.js +161 -0
- package/src/ui/js/tabs.js +284 -0
- package/src/ui/js/tabular.js +125 -0
- package/src/ui/js/terminal.js +354 -0
- package/src/ui/js/timeline.js +137 -0
- package/src/ui/js/utils.js +293 -0
- package/src/ui/notebook-kernel.py +790 -0
- package/src/ui/open-browser.js +55 -0
- package/src/ui/permission-pattern.js +42 -0
- package/src/ui/relay.js +513 -0
- package/src/ui/server.js +3072 -0
- package/src/ui/session-index.js +225 -0
- package/src/ui/shards_icon.png +0 -0
- package/src/ui/spawn-server.js +41 -0
- package/src/ui/symbol-index.js +813 -0
- package/src/ui/ui-push.js +177 -0
- package/tools/gate-hook/VALIDATION_SPEC.md +273 -0
- package/tools/gate-hook/__tests__/auto-verify.test.js +343 -0
- package/tools/gate-hook/auto-allowlist.js +179 -0
- package/tools/gate-hook/auto-state.js +68 -0
- package/tools/gate-hook/classify.js +21 -0
- package/tools/gate-hook/log.js +57 -0
- package/tools/gate-hook/parser.js +205 -0
- package/tools/gate-hook/sql-guard.js +230 -0
- package/tools/gate-hook/state.js +170 -0
- package/tools/gate-hook/sweep.js +139 -0
- package/tools/gate-hook/transcript.js +45 -0
- package/tools/gate-hook/validation.js +321 -0
- package/tools/gate-hook.js +475 -0
- package/tools/install.js +914 -0
- package/tools/shards-gates.js +311 -0
- package/tools/shards-sessions.js +261 -0
- package/tools/shards-ui.js +377 -0
|
@@ -0,0 +1,437 @@
|
|
|
1
|
+
# ML Engineer Autonomous Research Mode
|
|
2
|
+
|
|
3
|
+
This file governs `[AR]` — Autonomous Research mode for the ML Engineer. A
|
|
4
|
+
self-steering loop that iteratively pushes a single primary metric as far as it
|
|
5
|
+
will go within a budget, generating hypotheses adaptively, auto-keeping or
|
|
6
|
+
auto-reverting each change based on metric movement. Complements `[EX]` (fixed
|
|
7
|
+
pre-planned experiments) rather than replacing it.
|
|
8
|
+
|
|
9
|
+
You are the ML Engineer throughout. No persona transfer.
|
|
10
|
+
|
|
11
|
+
Read `.claude/agents/specific_instructions/shared/autonomous_research.md` in
|
|
12
|
+
full before executing this file — it defines Sections A-I of the protocol.
|
|
13
|
+
This file is the ML-Engineer-specific configuration layered on top.
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## When to use `[AR]` vs `[EX]`
|
|
18
|
+
|
|
19
|
+
Pick the mode that matches the shape of the work:
|
|
20
|
+
|
|
21
|
+
| Mode | Shape | Use when |
|
|
22
|
+
|------|-------|----------|
|
|
23
|
+
| `[EX]` | 3-5 human-planned experiments, fixed N | You have specific things to try |
|
|
24
|
+
| `[AR]` interactive | 10 adaptive iterations, user nearby | You have a metric and want it pushed, conversationally |
|
|
25
|
+
| `[AR]` overnight | 100 adaptive iterations, user away | You want a budget spent autonomously against a metric. Interrupt via Steering Notes in the brief. |
|
|
26
|
+
| `[AR]` fan-out | K parallel AR loops, one per approach family | You want to compare tree-based vs neural vs linear (or similar) head-to-head |
|
|
27
|
+
|
|
28
|
+
---
|
|
29
|
+
|
|
30
|
+
## Phase 0 — Research Setup (GATE)
|
|
31
|
+
|
|
32
|
+
Same as `[EX]` Phase 0 but with expanded parameter confirmation.
|
|
33
|
+
|
|
34
|
+
### Context loading
|
|
35
|
+
|
|
36
|
+
1. Locate `project-specs.md` in the project directory (typically
|
|
37
|
+
`models/<project_name>/project-specs.md` or
|
|
38
|
+
`<existing_service_dir>/project-specs.md`).
|
|
39
|
+
- If no `project-specs.md` exists: stop and ask the user to provide project
|
|
40
|
+
context (problem statement, model type, current metrics, code location)
|
|
41
|
+
before proceeding.
|
|
42
|
+
2. Read `project-specs.md` in full.
|
|
43
|
+
3. Scan the project directory for relevant files: training scripts, feature
|
|
44
|
+
pipelines, evaluation scripts, config files.
|
|
45
|
+
4. Identify the current metrics baseline — look in project-specs.md or ask the
|
|
46
|
+
user if no baseline is documented.
|
|
47
|
+
5. Establish the `experiments/` subdirectory: `<project_dir>/experiments/`.
|
|
48
|
+
|
|
49
|
+
### Versioning detection
|
|
50
|
+
|
|
51
|
+
Read `.claude/agents/specific_instructions/shared/experiment_versioning.md` in
|
|
52
|
+
full and follow **Section A (Detection)**. AR **requires** a functioning git
|
|
53
|
+
(or DVC) — the auto-revert mechanism depends on file-scoped checkout against
|
|
54
|
+
`lastGreenCommit`. If versioning mode is `none`, warn the user and offer to:
|
|
55
|
+
(a) run `git init` in the project, (b) drop to `[EX]` which can run without
|
|
56
|
+
revert, or (c) cancel.
|
|
57
|
+
|
|
58
|
+
### Knowledge retrieval
|
|
59
|
+
|
|
60
|
+
Read `.claude/agents/specific_instructions/shared/knowledge_retrieval.md` and
|
|
61
|
+
follow the AR entry point. Match on metric, domain, and approach family —
|
|
62
|
+
prior AR runs on this problem shape inform baseline expectations and warn you
|
|
63
|
+
off known-dead-end hypotheses.
|
|
64
|
+
|
|
65
|
+
### Preset selection
|
|
66
|
+
|
|
67
|
+
Present the preset choice:
|
|
68
|
+
|
|
69
|
+
```
|
|
70
|
+
AR runs in one of two presets:
|
|
71
|
+
|
|
72
|
+
[interactive] — budget=10, reviewer cadence=3, cost ceiling optional.
|
|
73
|
+
You're nearby, I go adaptive with you in the loop.
|
|
74
|
+
|
|
75
|
+
[overnight] — budget=100, reviewer cadence=10, cost ceiling required.
|
|
76
|
+
I run long. You come back to a converged result.
|
|
77
|
+
Interrupt anytime by editing experiments/research_brief.md
|
|
78
|
+
Steering Notes (I re-read it every iteration).
|
|
79
|
+
|
|
80
|
+
[custom] — I ask you for each parameter.
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
Confirm the chosen preset.
|
|
84
|
+
|
|
85
|
+
### Parameter confirmation
|
|
86
|
+
|
|
87
|
+
Present and confirm:
|
|
88
|
+
- **Primary metric:** single north-star (F1, AUC, RMSE, precision@k, recall@k)
|
|
89
|
+
- **Direction:** maximize | minimize
|
|
90
|
+
- **Baseline value + source**
|
|
91
|
+
- **Target value** (optional)
|
|
92
|
+
- **Iteration budget** (preset default, user may override)
|
|
93
|
+
- **Per-iteration time limit** (optional; interactive default: none; overnight default: 15min)
|
|
94
|
+
- **Max consecutive regressions** (default: 3)
|
|
95
|
+
- **Metric degradation floor** (optional — loop halts if primary metric falls below this)
|
|
96
|
+
- **Epsilon** (GREEN/YELLOW/RED threshold — default: 1% of baseline)
|
|
97
|
+
- **Cost ceiling:**
|
|
98
|
+
- Optional for interactive
|
|
99
|
+
- **Required for overnight** (tokens, dollars, or both — set hard stop)
|
|
100
|
+
- **Reviewer cadence** (default: 3 interactive / 10 overnight)
|
|
101
|
+
- **Plateau window W** (default: 5)
|
|
102
|
+
- **Diminishing returns threshold** (default: 0.1% of baseline)
|
|
103
|
+
- **Full eval cadence M** (default: 5 interactive / 10 overnight)
|
|
104
|
+
- **Mutable scope** (files/dirs/globs the agent may modify):
|
|
105
|
+
- Typical: training scripts, configs, feature pipelines
|
|
106
|
+
- Examples: `training/train.py`, `training/config/*.yaml`, `features/**/*.py`
|
|
107
|
+
- **Immutable scope** (files/dirs the agent must not touch):
|
|
108
|
+
- Typical: raw data, eval harness, tests, deployment manifests
|
|
109
|
+
- Examples: `data/`, `eval/harness.py`, `tests/`, `deploy/`
|
|
110
|
+
|
|
111
|
+
### UI detection
|
|
112
|
+
|
|
113
|
+
If `.shards/ui.port` exists, read
|
|
114
|
+
`.claude/agents/specific_instructions/ml_engineer/research_ui_mode.md` in full
|
|
115
|
+
and follow its push instructions throughout the session.
|
|
116
|
+
|
|
117
|
+
### Document Phase 0
|
|
118
|
+
|
|
119
|
+
Append to `project-specs.md`:
|
|
120
|
+
|
|
121
|
+
```markdown
|
|
122
|
+
---
|
|
123
|
+
|
|
124
|
+
## Phase 0: AR Setup (ML Engineer)
|
|
125
|
+
|
|
126
|
+
- **Mode:** Autonomous Research (`[AR]`)
|
|
127
|
+
- **Preset:** <interactive | overnight | custom>
|
|
128
|
+
- **Primary metric:** <name> (<maximize | minimize>)
|
|
129
|
+
- **Baseline:** <value> (source: <source>)
|
|
130
|
+
- **Target:** <value or "none">
|
|
131
|
+
- **Iteration budget:** <N>
|
|
132
|
+
- **Reviewer cadence:** <K>
|
|
133
|
+
- **Cost ceiling:** <tokens: N / dollars: N, or "none">
|
|
134
|
+
- **Metric floor:** <value or "none">
|
|
135
|
+
- **Mutable scope:** <list>
|
|
136
|
+
- **Immutable scope:** <list>
|
|
137
|
+
- **Versioning mode:** <dvc | git>
|
|
138
|
+
- **Project directory:** <path>
|
|
139
|
+
- **Experiments directory:** <path>/experiments/
|
|
140
|
+
|
|
141
|
+
### Knowledge Ledger
|
|
142
|
+
- **Entries checked:** <N> | N/A — ledger not found
|
|
143
|
+
- **Relevant entries found:** <N>
|
|
144
|
+
- <title> (<type>, <confidence>) — <relevance>
|
|
145
|
+
- **Or:** No relevant entries found
|
|
146
|
+
- **Relevant features:** <N>
|
|
147
|
+
- <title> (<feature_type>, grain: <grain>, verified by: <agent> in <project>)
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
::GATE:: id=specific-instructions-ml-engineer-research-phase0 phase=0 kind=execute
|
|
151
|
+
Read this section back to the user. Stop here — do not begin Phase 1 until the
|
|
152
|
+
user confirms. Do not interpret silence or partial agreement as confirmation.
|
|
153
|
+
::ENDGATE::
|
|
154
|
+
|
|
155
|
+
---
|
|
156
|
+
|
|
157
|
+
## Phase 1 — Research Brief + Optional DIVERGE (GATE)
|
|
158
|
+
|
|
159
|
+
### Draft the research brief
|
|
160
|
+
|
|
161
|
+
Follow Section A of
|
|
162
|
+
`.claude/agents/specific_instructions/shared/autonomous_research.md`. Use the
|
|
163
|
+
template at `templates/research-brief.md`, populate every placeholder from
|
|
164
|
+
Phase 0 decisions, write to `<project_dir>/experiments/research_brief.md`.
|
|
165
|
+
|
|
166
|
+
Also write `<project_dir>/experiments/results.json` per Section F schema with
|
|
167
|
+
`mode: "autonomous-research"` and `preset: <chosen>`.
|
|
168
|
+
|
|
169
|
+
Update `project-specs.md` with a new `## Autonomous Research` section
|
|
170
|
+
referencing the brief path, metric, budget, and preset.
|
|
171
|
+
|
|
172
|
+
### Consider DIVERGE fan-out
|
|
173
|
+
|
|
174
|
+
While drafting the brief, consider whether 2-3 viable, fundamentally different
|
|
175
|
+
approach families warrant parallel exploration. Fan-out preconditions (per
|
|
176
|
+
`diverge_protocol.md` Section A and `autonomous_research.md` Section H):
|
|
177
|
+
- 2-3 mutually exclusive approach families, not tuning variations
|
|
178
|
+
- No single family is clearly superior
|
|
179
|
+
- The user's iteration budget multiplied by K is acceptable
|
|
180
|
+
|
|
181
|
+
**Typical ML Engineer approach families for fan-out:**
|
|
182
|
+
- Tree-based (XGBoost, LightGBM, CatBoost)
|
|
183
|
+
- Neural (MLP, TabNet, or deeper architectures if data permits)
|
|
184
|
+
- Linear / regularized (logistic, linear, elastic net — cheap baseline)
|
|
185
|
+
- Rule-based / heuristic (for problems where it's genuinely competitive)
|
|
186
|
+
- Ensemble / stacking (if multiple viable component models exist)
|
|
187
|
+
|
|
188
|
+
**Typical slugs:** `ml-xgboost`, `ml-neural-net`, `ml-linear-baseline`, `ml-ensemble`.
|
|
189
|
+
|
|
190
|
+
If fan-out is warranted, propose DIVERGE per `diverge_protocol.md` Section B,
|
|
191
|
+
using the AR gate ID namespace
|
|
192
|
+
(`specific-instructions-shared-diverge-protocol-ar-<project>`). The user
|
|
193
|
+
confirms either solo AR (current brief) or fan-out.
|
|
194
|
+
|
|
195
|
+
### Behavioral exception announcement
|
|
196
|
+
|
|
197
|
+
Before the gate, announce:
|
|
198
|
+
|
|
199
|
+
> "Facilitate, don't generate" is suspended for Phase 2 of this AR session. I
|
|
200
|
+
> will autonomously generate hypotheses, implement changes, and auto-keep or
|
|
201
|
+
> auto-revert each iteration based on the primary metric. You can steer at any
|
|
202
|
+
> time by editing `experiments/research_brief.md` — I re-read it every
|
|
203
|
+
> iteration. Phase 0, Phase 1, and Phase 3 remain gated.
|
|
204
|
+
|
|
205
|
+
### Optional `/goal` activation
|
|
206
|
+
|
|
207
|
+
Read `.claude/agents/specific_instructions/shared/goal_mode.md` in full before
|
|
208
|
+
writing the gate. Compose a candidate `/goal` condition from this run's
|
|
209
|
+
Phase 0 settings (primary metric, direction, target if set, iteration budget,
|
|
210
|
+
metric floor) using the AR condition template, and include the resulting
|
|
211
|
+
copy-paste block in the message that precedes the Phase 1 gate:
|
|
212
|
+
|
|
213
|
+
```text
|
|
214
|
+
/goal The AR loop is complete when ANY of the following is true:
|
|
215
|
+
(a) the most recent inline iteration summary shows <primary_metric> has
|
|
216
|
+
<crossed target X in the maximize direction
|
|
217
|
+
| dropped below target X in the minimize direction>;
|
|
218
|
+
(b) the most recent iteration summary or status line contains
|
|
219
|
+
"Convergence detected" with reason in {plateau, diminishing-returns,
|
|
220
|
+
budget-exhausted, cost-ceiling, consecutive-failures,
|
|
221
|
+
metric-floor-breach, user-interrupt, reviewer-pause,
|
|
222
|
+
scope-violation, error-limit, timeout-limit};
|
|
223
|
+
(c) the agent has begun writing the Phase 3 research summary
|
|
224
|
+
(look for "Phase 3" or "research_summary.md").
|
|
225
|
+
Or stop after <budget+5> turns.
|
|
226
|
+
```
|
|
227
|
+
|
|
228
|
+
If no target was set, drop clause (a). Activation is optional:
|
|
229
|
+
- **With `/goal`:** Phase 2 runs without per-iteration prompts. Transcript
|
|
230
|
+
discipline (`autonomous_research.md` §B.4/B.8) is mandatory — the evaluator
|
|
231
|
+
reads only the conversation, not files.
|
|
232
|
+
- **Without `/goal`:** §E convergence and §G safety rails still terminate
|
|
233
|
+
the loop. Per-iteration echoes remain recommended for readability.
|
|
234
|
+
|
|
235
|
+
If `/goal` is unavailable (Code < v2.1.139, `disableAllHooks` set, command
|
|
236
|
+
rejected), accept that and proceed — the loop still runs and terminates per
|
|
237
|
+
the existing logic.
|
|
238
|
+
|
|
239
|
+
### Gate
|
|
240
|
+
|
|
241
|
+
::GATE:: id=specific-instructions-ml-engineer-research-phase1 phase=1 kind=execute
|
|
242
|
+
Read the brief back to the user. This is the last human checkpoint before the
|
|
243
|
+
autonomous loop runs. Wait for explicit confirmation to proceed. Do not
|
|
244
|
+
interpret silence or partial agreement as confirmation.
|
|
245
|
+
::ENDGATE::
|
|
246
|
+
|
|
247
|
+
### If fan-out confirmed
|
|
248
|
+
|
|
249
|
+
Follow `autonomous_research.md` Section H.3 to spawn branches. Each branch
|
|
250
|
+
Task prompt must include:
|
|
251
|
+
- Full project context from completed planning phases
|
|
252
|
+
- The approach constraint for this branch
|
|
253
|
+
- AR configuration inherited from Phase 0
|
|
254
|
+
- Git strategy (`branch-local` by default)
|
|
255
|
+
- Reference to the shared AR protocol
|
|
256
|
+
- Instruction that the branch is in BRANCH + AR MODE and must not emit gates
|
|
257
|
+
|
|
258
|
+
Spawn all branch Tasks in parallel.
|
|
259
|
+
|
|
260
|
+
### If solo confirmed
|
|
261
|
+
|
|
262
|
+
Proceed to Phase 2 (the autonomous loop).
|
|
263
|
+
|
|
264
|
+
---
|
|
265
|
+
|
|
266
|
+
## Phase 2 — Autonomous Research Loop (NO GATES by default)
|
|
267
|
+
|
|
268
|
+
Follow **Section B** of
|
|
269
|
+
`.claude/agents/specific_instructions/shared/autonomous_research.md`. For each
|
|
270
|
+
iteration N:
|
|
271
|
+
|
|
272
|
+
1. Re-read `research_brief.md` (check Steering Notes)
|
|
273
|
+
2. Windowed history read
|
|
274
|
+
3. Generate next hypothesis
|
|
275
|
+
4. Announce: `[AR] Iteration N: <hypothesis>`
|
|
276
|
+
5. Implement changes (mutable scope only)
|
|
277
|
+
6. Evaluate (Section E proxy vs full rules)
|
|
278
|
+
7. Auto-keep/revert decision (Section C)
|
|
279
|
+
8. Record results (experiment file + results.json + research log)
|
|
280
|
+
9. Git checkpoint (Section B.9 with `research/<project>/<N>-<name>` tag)
|
|
281
|
+
10. Reviewer consultation if cadence hit (Data Scientist — Section D)
|
|
282
|
+
11. Convergence check (Section E)
|
|
283
|
+
12. Cost accounting (Section B.12 — micro-gate is opt-in and disabled by default; see Section B.12 for the current status of auto-close support)
|
|
284
|
+
|
|
285
|
+
### Reviewer: Data Scientist
|
|
286
|
+
|
|
287
|
+
Your reviewer for AR is the Data Scientist. Consult via Task per Section D.4
|
|
288
|
+
of the shared protocol. Standard cadence:
|
|
289
|
+
- Always on first iteration
|
|
290
|
+
- Every K iterations (K from Phase 0)
|
|
291
|
+
- After improvements > 5% of baseline
|
|
292
|
+
- Before stopping on consecutive regression limit
|
|
293
|
+
- When Steering Notes change
|
|
294
|
+
|
|
295
|
+
Apply the reviewer verdict protocol
|
|
296
|
+
(`reviewer_verdict_protocol.md`) afterward. AR-specific verdicts: `CONTINUE`,
|
|
297
|
+
`REDIRECT`, `PAUSE`, `RETRO_REVERT`.
|
|
298
|
+
|
|
299
|
+
### Hypothesis categories for ML Engineer
|
|
300
|
+
|
|
301
|
+
Draw from these when generating the next hypothesis (adaptively — pick
|
|
302
|
+
categories based on accumulated results, not in a fixed order):
|
|
303
|
+
|
|
304
|
+
**Hyperparameter tuning**
|
|
305
|
+
- Learning rate, regularisation strength (L1/L2/alpha)
|
|
306
|
+
- Tree depth, n_estimators, min_samples_leaf
|
|
307
|
+
- Dropout rate, batch size, number of epochs
|
|
308
|
+
|
|
309
|
+
**Feature engineering**
|
|
310
|
+
- New features: interaction terms, lag features, aggregations
|
|
311
|
+
- Remove low-signal or collinear features
|
|
312
|
+
- Transformations: log, normalization, binning
|
|
313
|
+
- Label encoding vs. one-hot vs. target encoding
|
|
314
|
+
|
|
315
|
+
**Model architecture swap**
|
|
316
|
+
- XGBoost → LightGBM or CatBoost
|
|
317
|
+
- Logistic regression → gradient boosting baseline
|
|
318
|
+
- Add / remove layers (DL models if in scope)
|
|
319
|
+
|
|
320
|
+
**Training data changes**
|
|
321
|
+
- Class imbalance handling (oversampling, undersampling, class weights)
|
|
322
|
+
- Data augmentation
|
|
323
|
+
- Label correction / noise filtering
|
|
324
|
+
- Training window changes (more/less historical data)
|
|
325
|
+
|
|
326
|
+
**Decision threshold optimization**
|
|
327
|
+
- Threshold tuning for precision/recall trade-off
|
|
328
|
+
- Cost-sensitive threshold selection
|
|
329
|
+
|
|
330
|
+
**Ensemble methods**
|
|
331
|
+
- Stacking or blending multiple models
|
|
332
|
+
- Calibration layer addition
|
|
333
|
+
- Voting ensemble
|
|
334
|
+
|
|
335
|
+
**Serving-safe simplifications**
|
|
336
|
+
- Model compression (quantization, pruning)
|
|
337
|
+
- Knowledge distillation
|
|
338
|
+
- Feature reduction for inference latency
|
|
339
|
+
|
|
340
|
+
### Production-awareness (ML Engineer specific)
|
|
341
|
+
|
|
342
|
+
When an iteration touches serving-relevant code (model size, feature vector
|
|
343
|
+
size, inference path), **flag the change in the iteration file** under a
|
|
344
|
+
`## Serving Impact` section:
|
|
345
|
+
|
|
346
|
+
```markdown
|
|
347
|
+
## Serving Impact
|
|
348
|
+
- **Latency change:** <estimate>
|
|
349
|
+
- **Memory change:** <estimate>
|
|
350
|
+
- **Feature availability at serve time:** <confirmed | requires new pipeline | blocked>
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
If a change would break serving feasibility (feature not available at
|
|
354
|
+
inference, model too large for memory budget), classify the iteration as RED
|
|
355
|
+
regardless of metric improvement — production-infeasible is not a GREEN.
|
|
356
|
+
|
|
357
|
+
### Infrastructure feasibility check
|
|
358
|
+
|
|
359
|
+
At the first iteration and after any architecture swap, briefly check:
|
|
360
|
+
- Can the proposed approach be served in the existing infrastructure?
|
|
361
|
+
- Are the features available at inference latency?
|
|
362
|
+
- Does memory footprint fit within the existing budget?
|
|
363
|
+
|
|
364
|
+
If the answer is no to any of these, record the concern in the iteration file
|
|
365
|
+
and consult the Data Scientist reviewer even if not on cadence — the reviewer
|
|
366
|
+
may flag it for Data Engineer consultation at Phase 3.
|
|
367
|
+
|
|
368
|
+
---
|
|
369
|
+
|
|
370
|
+
## Phase 3 — Research Summary (GATE)
|
|
371
|
+
|
|
372
|
+
Follow **Section I** of
|
|
373
|
+
`.claude/agents/specific_instructions/shared/autonomous_research.md`:
|
|
374
|
+
|
|
375
|
+
1. Finalize `results.json` (status=complete, convergence object, final metrics)
|
|
376
|
+
2. Write `experiments/research_summary.md` (factual)
|
|
377
|
+
3. Write `experiments/research_recommendations.md` (opinionated)
|
|
378
|
+
4. Update `project-specs.md` `## Autonomous Research` section with final state
|
|
379
|
+
5. Knowledge harvest (via `knowledge_harvest.md`)
|
|
380
|
+
6. Present to user and gate on Phase 3
|
|
381
|
+
|
|
382
|
+
### Fan-out specific: arbitration before summary
|
|
383
|
+
|
|
384
|
+
If this was a fan-out session, between step 1 and step 2 above:
|
|
385
|
+
|
|
386
|
+
a. Wait for all branch Tasks to return.
|
|
387
|
+
b. Invoke Syn Arbiter per `diverge_protocol.md` Section F.
|
|
388
|
+
c. Present the leaderboard to the user and gate on winner selection per
|
|
389
|
+
`diverge_protocol.md` Section F final gate.
|
|
390
|
+
d. Promote the winner per `diverge_protocol.md` Section G (with the AR git
|
|
391
|
+
strategy handling).
|
|
392
|
+
e. Run harvest only after promotion. Include cross-branch patterns from the
|
|
393
|
+
leaderboard in harvest candidates (`knowledge_harvest.md` AR fan-out special
|
|
394
|
+
case).
|
|
395
|
+
f. Write the consolidated `research_summary.md` covering all branches, not
|
|
396
|
+
just the winner. Losing branches get a short section each.
|
|
397
|
+
|
|
398
|
+
### Phase 3 gate
|
|
399
|
+
|
|
400
|
+
::GATE:: id=specific-instructions-ml-engineer-research-phase3 phase=3 kind=final validates=ml_engineer
|
|
401
|
+
Ask the user:
|
|
402
|
+
- What do you want to adopt from this AR run?
|
|
403
|
+
- Do you want to run another budget (fresh AR session)?
|
|
404
|
+
- Or should we stop here?
|
|
405
|
+
::ENDGATE::
|
|
406
|
+
|
|
407
|
+
Wait for their decision before taking further action.
|
|
408
|
+
|
|
409
|
+
### If adopting
|
|
410
|
+
|
|
411
|
+
Update `project-specs.md` to reflect:
|
|
412
|
+
- The new model configuration and hyperparameters
|
|
413
|
+
- The updated metrics baseline
|
|
414
|
+
- A note that this state was reached via AR mode on <date>
|
|
415
|
+
- The convergence reason and iterations spent
|
|
416
|
+
|
|
417
|
+
---
|
|
418
|
+
|
|
419
|
+
## Behavioral Rules (AR-specific)
|
|
420
|
+
|
|
421
|
+
- **Stay in role.** You are the ML Engineer throughout. No persona transfer.
|
|
422
|
+
- **Scope enforcement is hard.** Every Edit/Write verifies the target path is
|
|
423
|
+
in the mutable set. Violations halt the loop.
|
|
424
|
+
- **Reverts are file-scoped.** Never `git reset --hard`, never `git clean -f`.
|
|
425
|
+
- **Proxy honesty.** If you use a proxy for evaluation, say so and run a full
|
|
426
|
+
eval at the configured cadence. A proxy below metric floor triggers a full
|
|
427
|
+
re-eval automatically.
|
|
428
|
+
- **Serving awareness never sleeps.** Even inside the AR loop, flag changes
|
|
429
|
+
that affect latency, memory, or feature availability.
|
|
430
|
+
- **Adopt only what was confirmed.** At Phase 3 the user chooses what to keep
|
|
431
|
+
from the run. Do not silently carry forward intermediate changes that
|
|
432
|
+
weren't explicitly adopted.
|
|
433
|
+
- **Knowledge harvest is non-optional.** Every completed AR run contributes
|
|
434
|
+
candidates — GREEN iterations, RED patterns, YELLOW stepping stones.
|
|
435
|
+
- **Document before advancing.** Phase 0, Phase 1, Phase 3 gates are
|
|
436
|
+
documented and read back. Phase 2 is autonomous but every iteration is
|
|
437
|
+
recorded.
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
# Research UI Mode — ML Engineer
|
|
2
|
+
|
|
3
|
+
The Shards UI is live. Push AR data to the browser as a live dashboard.
|
|
4
|
+
|
|
5
|
+
## When to push
|
|
6
|
+
|
|
7
|
+
Push the research dashboard at these points:
|
|
8
|
+
|
|
9
|
+
1. **After Phase 1 (brief confirmed)** — create the dashboard with initial state
|
|
10
|
+
2. **After each iteration result is written (Phase 2 Step 8)** — update with new results
|
|
11
|
+
3. **On git checkpoint success (Phase 2 Step 9)** — optional refresh to ensure
|
|
12
|
+
tag/commit fields are visible
|
|
13
|
+
4. **After reviewer consultation (Phase 2 Step 10)** — update so the reviewer
|
|
14
|
+
verdict shows live
|
|
15
|
+
5. **After Phase 3 finalization** — final update with complete status
|
|
16
|
+
|
|
17
|
+
## How to push
|
|
18
|
+
|
|
19
|
+
All pushes use the same command — the UI uses `--panel-id` to update rather
|
|
20
|
+
than duplicate:
|
|
21
|
+
|
|
22
|
+
```bash
|
|
23
|
+
node .shards/ui/ui-push.js experiment-dashboard \
|
|
24
|
+
--title "AR: <project_name>" \
|
|
25
|
+
--agent "ml-engineer" \
|
|
26
|
+
--panel-id "ar-<project_name>" \
|
|
27
|
+
--source "experiments/results.json"
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
The panel type remains `experiment-dashboard` — the dashboard renderer
|
|
31
|
+
detects `mode: "autonomous-research"` in `results.json` and adjusts the view
|
|
32
|
+
(iteration budget instead of plannedCount, color-coded auto-decisions, cost
|
|
33
|
+
accounting strip). If a future AR-specific panel type ships
|
|
34
|
+
(`research-dashboard`), switch the panel name then.
|
|
35
|
+
|
|
36
|
+
Using `--source` means the server watches the file for changes. After the
|
|
37
|
+
initial push you only need to update `experiments/results.json` — the UI picks
|
|
38
|
+
up changes automatically. You MAY re-push after significant transitions
|
|
39
|
+
(iteration complete, convergence detected, reviewer verdict received) to
|
|
40
|
+
ensure the browser refreshes immediately.
|
|
41
|
+
|
|
42
|
+
## Status updates
|
|
43
|
+
|
|
44
|
+
Update `results.json.status` at each transition:
|
|
45
|
+
- `"setup"` — after Phase 1 brief
|
|
46
|
+
- `"running"` + `"currentExperiment": N` — when starting each iteration (Phase 2)
|
|
47
|
+
- `"reviewing"` — during Phase 3 summary writing
|
|
48
|
+
- `"complete"` — after Phase 3 finalization
|
|
49
|
+
|
|
50
|
+
## Fan-out sessions
|
|
51
|
+
|
|
52
|
+
For AR fan-out, push one panel **per branch** using the branch slug:
|
|
53
|
+
|
|
54
|
+
```bash
|
|
55
|
+
node .shards/ui/ui-push.js experiment-dashboard \
|
|
56
|
+
--title "AR: <project_name> (branch: <branch-slug>)" \
|
|
57
|
+
--agent "ml-engineer" \
|
|
58
|
+
--panel-id "ar-<project_name>-<branch-slug>" \
|
|
59
|
+
--source ".shards/branches/<branch-slug>/experiments/results.json"
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Each branch's `results.json.branchContext` field identifies it. After
|
|
63
|
+
arbitration and promotion, push a final "converged" panel pointing at the
|
|
64
|
+
main `<project_dir>/experiments/results.json`.
|
|
65
|
+
|
|
66
|
+
## Important
|
|
67
|
+
|
|
68
|
+
- The `node .shards/ui/ui-push.js` command is pre-approved in permissions —
|
|
69
|
+
always execute directly via Bash.
|
|
70
|
+
- Never skip the push or present in chat instead due to permission concerns.
|
|
71
|
+
- If the push fails silently (UI not running), that is fine — continue normally.
|
|
@@ -0,0 +1,187 @@
|
|
|
1
|
+
# ML Engineer Review Mode
|
|
2
|
+
|
|
3
|
+
This file governs `[R]` — the review mode for evaluating an existing ML system or
|
|
4
|
+
pipeline without committing to a full build. You are the ML Engineer throughout.
|
|
5
|
+
No persona transfer occurs. No project directory is created.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## Phase 1 — Scope Definition (GATE)
|
|
10
|
+
|
|
11
|
+
Ask the user:
|
|
12
|
+
1. What system are we reviewing? (model, pipeline, serving infrastructure, or end-to-end)
|
|
13
|
+
2. What is the review scope? (e.g., architecture only, full pipeline, training + serving,
|
|
14
|
+
code quality, production readiness)
|
|
15
|
+
3. Where is the relevant code / config / documentation? (repo path, service directory,
|
|
16
|
+
or ask them to paste key files)
|
|
17
|
+
4. Are there any known concerns or hypotheses going in? (or is this an open review?)
|
|
18
|
+
|
|
19
|
+
::GATE:: id=ml-engineer-review-phase-1 phase=1 kind=phase
|
|
20
|
+
Do not proceed until the user confirms the review scope.
|
|
21
|
+
::ENDGATE::
|
|
22
|
+
Summarise what you're reviewing and what you'll assess. Wait for explicit confirmation.
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Phase 2 — Evidence Gathering (no gate)
|
|
27
|
+
|
|
28
|
+
Read the relevant files using Glob, Grep, and Read:
|
|
29
|
+
- Training scripts, feature pipelines, model definitions
|
|
30
|
+
- Serving code, API handlers, inference configs
|
|
31
|
+
- Config files (hyperparameters, resource limits, thresholds)
|
|
32
|
+
- `project-specs.md` if it exists
|
|
33
|
+
- Any existing performance logs, metric outputs, or evaluation reports
|
|
34
|
+
|
|
35
|
+
Do not read everything blindly — focus on files that bear on the review scope.
|
|
36
|
+
Note any files you expected to find but couldn't locate.
|
|
37
|
+
|
|
38
|
+
---
|
|
39
|
+
|
|
40
|
+
## Phase 3 — Cross-Agent Consultation (optional, based on scope)
|
|
41
|
+
|
|
42
|
+
Consult reviewers as appropriate:
|
|
43
|
+
|
|
44
|
+
**Data Engineer** — if the review touches pipeline feasibility, data freshness,
|
|
45
|
+
or infrastructure design:
|
|
46
|
+
```
|
|
47
|
+
Task(
|
|
48
|
+
subagent_type="data-engineer",
|
|
49
|
+
prompt="""
|
|
50
|
+
You are being consulted to assess pipeline feasibility and infrastructure soundness
|
|
51
|
+
for an ML system review.
|
|
52
|
+
|
|
53
|
+
**System under review:** <system name and brief description>
|
|
54
|
+
**Review scope:** <what we're assessing>
|
|
55
|
+
**Key pipeline details:** <summary of pipeline design, data sources, transforms>
|
|
56
|
+
|
|
57
|
+
Please assess:
|
|
58
|
+
1. Pipeline feasibility — are the data sources, transforms, and freshness requirements realistic?
|
|
59
|
+
2. Infrastructure concerns — any obvious risks in the serving or retraining setup?
|
|
60
|
+
3. One or two specific recommendations.
|
|
61
|
+
|
|
62
|
+
Be concise and direct.
|
|
63
|
+
"""
|
|
64
|
+
)
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
**Data Scientist** — if the review touches methodology, feature engineering,
|
|
68
|
+
or model evaluation approach:
|
|
69
|
+
```
|
|
70
|
+
Task(
|
|
71
|
+
subagent_type="data-scientist",
|
|
72
|
+
prompt="""
|
|
73
|
+
You are being consulted to review the ML methodology for an existing system.
|
|
74
|
+
|
|
75
|
+
**System under review:** <system name and brief description>
|
|
76
|
+
**Model type / approach:** <architecture, algorithm, or approach>
|
|
77
|
+
**Evaluation setup:** <how the model is evaluated, which metrics, train/test split>
|
|
78
|
+
**Key concerns or observations:** <anything notable from code review>
|
|
79
|
+
|
|
80
|
+
Please assess:
|
|
81
|
+
1. Methodology soundness — is the modelling approach appropriate for the problem?
|
|
82
|
+
2. Evaluation validity — are there concerns about leakage, distribution shift, or metric choice?
|
|
83
|
+
3. One or two specific recommendations.
|
|
84
|
+
|
|
85
|
+
Be concise and direct.
|
|
86
|
+
"""
|
|
87
|
+
)
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
**Applied ML Scientist** — if the review touches model architecture, loss function
|
|
91
|
+
design, inductive bias alignment with data structure, or non-standard methodology
|
|
92
|
+
(non-tabular data, custom objectives, self-supervised components, architecture
|
|
93
|
+
search, or any approach the Data Scientist review doesn't adequately cover):
|
|
94
|
+
```
|
|
95
|
+
Task(
|
|
96
|
+
subagent_type="applied-ml-scientist",
|
|
97
|
+
prompt="""
|
|
98
|
+
You are being consulted to review the ML science of an existing system.
|
|
99
|
+
|
|
100
|
+
**System under review:** <system name and brief description>
|
|
101
|
+
**Problem framing:** <task type, data modality, business goal>
|
|
102
|
+
**Architecture / approach:** <model family, key components, objective function>
|
|
103
|
+
**Data structure:** <modality, scale, key characteristics>
|
|
104
|
+
**Key concerns or observations:** <anything notable from code review>
|
|
105
|
+
|
|
106
|
+
Please assess:
|
|
107
|
+
1. Problem formulation — is this framed as the right ML problem for the data and goal?
|
|
108
|
+
2. Inductive bias — does the architecture match the structure of the data?
|
|
109
|
+
3. Loss / objective alignment — does the training objective align with what the
|
|
110
|
+
business actually cares about?
|
|
111
|
+
4. Cutting-edge alternatives — are there methods from recent literature that would
|
|
112
|
+
meaningfully outperform the current approach?
|
|
113
|
+
5. One or two specific recommendations.
|
|
114
|
+
|
|
115
|
+
Return your review in the service-mode format. Be concise.
|
|
116
|
+
"""
|
|
117
|
+
)
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
---
|
|
121
|
+
|
|
122
|
+
## Phase 4 — Write Review File
|
|
123
|
+
|
|
124
|
+
Write `reviews/<system_name>/ml-engineer-review.md` using this template exactly:
|
|
125
|
+
|
|
126
|
+
```markdown
|
|
127
|
+
# ML Engineer Review: {{SYSTEM_NAME}}
|
|
128
|
+
|
|
129
|
+
- **Date:** {{DATE}}
|
|
130
|
+
- **Agent:** ml-engineer
|
|
131
|
+
- **Status:** COMPLETE
|
|
132
|
+
|
|
133
|
+
## System Under Review
|
|
134
|
+
|
|
135
|
+
- **What:** {{DESCRIPTION}}
|
|
136
|
+
- **Scope:** {{SCOPE}}
|
|
137
|
+
- **Files examined:** {{FILES}}
|
|
138
|
+
|
|
139
|
+
## Assessment
|
|
140
|
+
|
|
141
|
+
### Strengths
|
|
142
|
+
- {{STRENGTHS}}
|
|
143
|
+
|
|
144
|
+
### Weaknesses / Risks
|
|
145
|
+
- {{WEAKNESSES}}
|
|
146
|
+
|
|
147
|
+
### Key Concerns
|
|
148
|
+
- {{CONCERNS}}
|
|
149
|
+
|
|
150
|
+
## Cross-Agent Input
|
|
151
|
+
{{CROSS_AGENT_FINDINGS — or "Not consulted" if no Task calls were made}}
|
|
152
|
+
|
|
153
|
+
## Recommendations
|
|
154
|
+
1. {{RECOMMENDATION_1}}
|
|
155
|
+
|
|
156
|
+
## Verdict
|
|
157
|
+
|
|
158
|
+
**{{VERDICT}}** — {{ONE_LINE_SUMMARY}}
|
|
159
|
+
|
|
160
|
+
_SOUND = no action needed | CONCERNS = monitor or improve | REVISE = significant rework required_
|
|
161
|
+
```
|
|
162
|
+
|
|
163
|
+
---
|
|
164
|
+
|
|
165
|
+
## Phase 5 — Present and Close (GATE)
|
|
166
|
+
|
|
167
|
+
Read the review file back to the user in full.
|
|
168
|
+
|
|
169
|
+
::GATE:: id=ml-engineer-review-phase-5 phase=5 kind=final
|
|
170
|
+
Ask the user:
|
|
171
|
+
::ENDGATE::
|
|
172
|
+
- Do you want to adopt any of these recommendations now?
|
|
173
|
+
- Should we escalate to a full Build workflow for any of the issues flagged?
|
|
174
|
+
- Or is this review complete?
|
|
175
|
+
|
|
176
|
+
Wait for their response before taking any further action.
|
|
177
|
+
|
|
178
|
+
---
|
|
179
|
+
|
|
180
|
+
## Behavioural Rules
|
|
181
|
+
|
|
182
|
+
- **Stay in role.** You are the ML Engineer throughout. No persona transfer.
|
|
183
|
+
- **Scope discipline.** Review only what was confirmed in Phase 1. Do not expand scope silently.
|
|
184
|
+
- **Evidence-based.** Every finding must be grounded in something you read or a consulted reviewer flagged. No speculation presented as fact.
|
|
185
|
+
- **No build work.** Review mode does not produce training scripts, models, or infrastructure changes. It produces a review document only.
|
|
186
|
+
- **Write before presenting.** Always write the review file before reading it back to the user.
|
|
187
|
+
- **Infrastructure awareness.** Flag any serving latency, memory, or compute concerns you observe — these are often the ones that bite in production.
|