@proflandrigan/shards 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +475 -0
- package/package.json +37 -0
- package/src/agents/academic.md +276 -0
- package/src/agents/ai-engineer.md +377 -0
- package/src/agents/analytics-engineer.md +364 -0
- package/src/agents/applied-ml-scientist.md +410 -0
- package/src/agents/backend-engineer.md +255 -0
- package/src/agents/bi-engineer.md +333 -0
- package/src/agents/data-analyst.md +343 -0
- package/src/agents/data-engineer.md +260 -0
- package/src/agents/data-modeller.md +386 -0
- package/src/agents/data-scientist.md +366 -0
- package/src/agents/deep-learning-engineer.md +389 -0
- package/src/agents/ml-engineer.md +424 -0
- package/src/agents/mlops-engineer.md +339 -0
- package/src/agents/researcher.md +187 -0
- package/src/agents/specific_instructions/academic/critical_review.md +263 -0
- package/src/agents/specific_instructions/academic/report.md +113 -0
- package/src/agents/specific_instructions/ai_engineer/advise.md +162 -0
- package/src/agents/specific_instructions/ai_engineer/bi_engineer_handoff.md +86 -0
- package/src/agents/specific_instructions/ai_engineer/experiment.md +471 -0
- package/src/agents/specific_instructions/ai_engineer/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/ai_engineer/phases/index.md +45 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-1.md +55 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-2.md +86 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-3.md +96 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-4.md +138 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-5.md +157 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-6.md +196 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-7.md +313 -0
- package/src/agents/specific_instructions/ai_engineer/phases.md +1011 -0
- package/src/agents/specific_instructions/ai_engineer/prompt_lab.md +161 -0
- package/src/agents/specific_instructions/ai_engineer/prompt_lab_ui_mode.md +28 -0
- package/src/agents/specific_instructions/ai_engineer/research.md +393 -0
- package/src/agents/specific_instructions/ai_engineer/research_ui_mode.md +66 -0
- package/src/agents/specific_instructions/ai_engineer/review.md +159 -0
- package/src/agents/specific_instructions/ai_engineer/validation_checklist.md +182 -0
- package/src/agents/specific_instructions/analytics_engineer/advise.md +155 -0
- package/src/agents/specific_instructions/analytics_engineer/bi_engineer_handoff.md +91 -0
- package/src/agents/specific_instructions/analytics_engineer/data_analyst_handoff.md +84 -0
- package/src/agents/specific_instructions/analytics_engineer/deep_phases.md +818 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/index.md +24 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-1.md +77 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-2.md +106 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-3.md +93 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-4.md +79 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-5.md +61 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-6.md +45 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-7.md +235 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-8.md +221 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-1.md +47 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-2.md +78 -0
- package/src/agents/specific_instructions/analytics_engineer/quick_phases.md +112 -0
- package/src/agents/specific_instructions/analytics_engineer/review.md +167 -0
- package/src/agents/specific_instructions/analytics_engineer/service_mode.md +369 -0
- package/src/agents/specific_instructions/analytics_engineer/ui_mode.md +45 -0
- package/src/agents/specific_instructions/analytics_engineer/update.md +162 -0
- package/src/agents/specific_instructions/analytics_engineer/validation_checklist.md +121 -0
- package/src/agents/specific_instructions/applied_ml_scientist/advise.md +143 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/index.md +21 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-1.md +51 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-2.md +66 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-3.md +113 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-4.md +104 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-5.md +156 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases.md +428 -0
- package/src/agents/specific_instructions/applied_ml_scientist/research.md +379 -0
- package/src/agents/specific_instructions/applied_ml_scientist/review.md +142 -0
- package/src/agents/specific_instructions/applied_ml_scientist/validation_checklist.md +136 -0
- package/src/agents/specific_instructions/backend_engineer/clean.md +149 -0
- package/src/agents/specific_instructions/backend_engineer/review.md +91 -0
- package/src/agents/specific_instructions/backend_engineer/review_checklist.md +54 -0
- package/src/agents/specific_instructions/backend_engineer/service_mode.md +67 -0
- package/src/agents/specific_instructions/bi_engineer/advise.md +137 -0
- package/src/agents/specific_instructions/bi_engineer/data_analyst_handoff.md +77 -0
- package/src/agents/specific_instructions/bi_engineer/incoming_handoff.md +45 -0
- package/src/agents/specific_instructions/bi_engineer/phases/index.md +20 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-1.md +164 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-2.md +92 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-3.md +121 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-4.md +106 -0
- package/src/agents/specific_instructions/bi_engineer/phases.md +451 -0
- package/src/agents/specific_instructions/bi_engineer/review.md +166 -0
- package/src/agents/specific_instructions/bi_engineer/update.md +147 -0
- package/src/agents/specific_instructions/bi_engineer/validation_checklist.md +124 -0
- package/src/agents/specific_instructions/data_analyst/advise.md +138 -0
- package/src/agents/specific_instructions/data_analyst/explain.md +221 -0
- package/src/agents/specific_instructions/data_analyst/incoming_handoff.md +40 -0
- package/src/agents/specific_instructions/data_analyst/phases/index.md +20 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-1.md +159 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-2.md +112 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-3.md +265 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-4.md +100 -0
- package/src/agents/specific_instructions/data_analyst/phases.md +501 -0
- package/src/agents/specific_instructions/data_analyst/review.md +138 -0
- package/src/agents/specific_instructions/data_analyst/ui_mode.md +26 -0
- package/src/agents/specific_instructions/data_analyst/update.md +144 -0
- package/src/agents/specific_instructions/data_analyst/validation_checklist.md +95 -0
- package/src/agents/specific_instructions/data_engineer/advise.md +137 -0
- package/src/agents/specific_instructions/data_engineer/phases.md +466 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/index.md +23 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-1.md +49 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-2.md +93 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-3.md +55 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-4.md +48 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-5.md +40 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-6.md +102 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-7.md +87 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/phase-1.md +45 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/phase-2.md +54 -0
- package/src/agents/specific_instructions/data_engineer/review.md +135 -0
- package/src/agents/specific_instructions/data_engineer/validation_checklist.md +136 -0
- package/src/agents/specific_instructions/data_modeller/advise.md +137 -0
- package/src/agents/specific_instructions/data_modeller/phases.md +581 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/index.md +23 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-1.md +52 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-2.md +113 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-3.md +47 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-4.md +51 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-5.md +45 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-6.md +105 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-7.md +136 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/phase-1.md +47 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/phase-2.md +65 -0
- package/src/agents/specific_instructions/data_modeller/review.md +141 -0
- package/src/agents/specific_instructions/data_modeller/service_mode.md +218 -0
- package/src/agents/specific_instructions/data_modeller/validation_checklist.md +125 -0
- package/src/agents/specific_instructions/data_scientist/advise.md +158 -0
- package/src/agents/specific_instructions/data_scientist/bi_engineer_handoff.md +63 -0
- package/src/agents/specific_instructions/data_scientist/experiment.md +482 -0
- package/src/agents/specific_instructions/data_scientist/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/data_scientist/explain.md +247 -0
- package/src/agents/specific_instructions/data_scientist/greenfield_data.md +35 -0
- package/src/agents/specific_instructions/data_scientist/ml_engineer_handoff.md +52 -0
- package/src/agents/specific_instructions/data_scientist/notebook_walkthrough.md +76 -0
- package/src/agents/specific_instructions/data_scientist/phases/index.md +24 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-1.md +45 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-2.md +67 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-3.md +89 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-4.md +143 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-5.md +71 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-6.md +239 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-7.md +207 -0
- package/src/agents/specific_instructions/data_scientist/phases.md +651 -0
- package/src/agents/specific_instructions/data_scientist/research.md +345 -0
- package/src/agents/specific_instructions/data_scientist/research_ui_mode.md +52 -0
- package/src/agents/specific_instructions/data_scientist/review.md +136 -0
- package/src/agents/specific_instructions/data_scientist/service_mode.md +247 -0
- package/src/agents/specific_instructions/data_scientist/validation_checklist.md +183 -0
- package/src/agents/specific_instructions/deep_learning_engineer/advise.md +145 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/index.md +21 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-1.md +74 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-2.md +98 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-3.md +76 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-4.md +128 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-5.md +292 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases.md +567 -0
- package/src/agents/specific_instructions/deep_learning_engineer/research.md +389 -0
- package/src/agents/specific_instructions/deep_learning_engineer/review.md +155 -0
- package/src/agents/specific_instructions/deep_learning_engineer/validation_checklist.md +147 -0
- package/src/agents/specific_instructions/ml_engineer/advise.md +174 -0
- package/src/agents/specific_instructions/ml_engineer/bi_engineer_handoff.md +71 -0
- package/src/agents/specific_instructions/ml_engineer/experiment.md +474 -0
- package/src/agents/specific_instructions/ml_engineer/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/ml_engineer/notebook_walkthrough.md +75 -0
- package/src/agents/specific_instructions/ml_engineer/phases/index.md +25 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-1.md +49 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-2.md +75 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-3.md +124 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-4.md +279 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-5.md +160 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-6-5.md +170 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-6.md +295 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-7.md +337 -0
- package/src/agents/specific_instructions/ml_engineer/phases.md +1068 -0
- package/src/agents/specific_instructions/ml_engineer/research.md +437 -0
- package/src/agents/specific_instructions/ml_engineer/research_ui_mode.md +71 -0
- package/src/agents/specific_instructions/ml_engineer/review.md +187 -0
- package/src/agents/specific_instructions/ml_engineer/service_mode.md +273 -0
- package/src/agents/specific_instructions/ml_engineer/validation_checklist.md +185 -0
- package/src/agents/specific_instructions/mlops_engineer/advise.md +139 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/index.md +23 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-1.md +52 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-2.md +86 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-3.md +105 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-4.md +128 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-5.md +106 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-6.md +128 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-7.md +144 -0
- package/src/agents/specific_instructions/mlops_engineer/phases.md +671 -0
- package/src/agents/specific_instructions/mlops_engineer/review.md +164 -0
- package/src/agents/specific_instructions/mlops_engineer/service_mode.md +81 -0
- package/src/agents/specific_instructions/mlops_engineer/validation_checklist.md +151 -0
- package/src/agents/specific_instructions/researcher/critical_review.md +292 -0
- package/src/agents/specific_instructions/researcher/review_checklist.md +67 -0
- package/src/agents/specific_instructions/researcher/service_mode.md +224 -0
- package/src/agents/specific_instructions/shared/auto_verify_mode.md +141 -0
- package/src/agents/specific_instructions/shared/autonomous_research.md +1289 -0
- package/src/agents/specific_instructions/shared/behavioral_rules.md +36 -0
- package/src/agents/specific_instructions/shared/diverge_protocol.md +387 -0
- package/src/agents/specific_instructions/shared/engineering_guidelines.md +136 -0
- package/src/agents/specific_instructions/shared/experiment_versioning.md +184 -0
- package/src/agents/specific_instructions/shared/goal_mode.md +187 -0
- package/src/agents/specific_instructions/shared/incremental_testing.md +139 -0
- package/src/agents/specific_instructions/shared/intent_discovery.md +223 -0
- package/src/agents/specific_instructions/shared/join_path_protocol.md +168 -0
- package/src/agents/specific_instructions/shared/knowledge_checkpoint.md +83 -0
- package/src/agents/specific_instructions/shared/knowledge_harvest.md +220 -0
- package/src/agents/specific_instructions/shared/knowledge_retrieval.md +100 -0
- package/src/agents/specific_instructions/shared/notebook_walkthrough_protocol.md +367 -0
- package/src/agents/specific_instructions/shared/reviewer_verdict_protocol.md +74 -0
- package/src/agents/specific_instructions/shared/swarm_protocol.md +97 -0
- package/src/agents/specific_instructions/shared/validation_protocol.md +139 -0
- package/src/agents/specific_instructions/syn/arbiter.md +140 -0
- package/src/agents/specific_instructions/syn/brainstorm.md +550 -0
- package/src/agents/specific_instructions/syn/code_review.md +232 -0
- package/src/agents/specific_instructions/syn/diff.md +239 -0
- package/src/agents/specific_instructions/syn/final_review.md +65 -0
- package/src/agents/specific_instructions/syn/fixer.md +240 -0
- package/src/agents/specific_instructions/syn/free_form.md +130 -0
- package/src/agents/specific_instructions/syn/knowledge.md +468 -0
- package/src/agents/specific_instructions/syn/notebook_walkthrough.md +78 -0
- package/src/agents/specific_instructions/syn/panel_review.md +634 -0
- package/src/agents/specific_instructions/syn/pm.md +453 -0
- package/src/agents/specific_instructions/syn/pr_review.md +255 -0
- package/src/agents/specific_instructions/syn/slides.md +417 -0
- package/src/agents/syn.md +729 -0
- package/src/commands/academic.md +41 -0
- package/src/commands/ai-engineer.md +45 -0
- package/src/commands/analytics-engineer.md +48 -0
- package/src/commands/applied-ml-scientist.md +45 -0
- package/src/commands/backend-engineer.md +35 -0
- package/src/commands/bi-engineer.md +40 -0
- package/src/commands/brainstorm.md +24 -0
- package/src/commands/data-analyst.md +38 -0
- package/src/commands/data-engineer.md +37 -0
- package/src/commands/data-modeller.md +38 -0
- package/src/commands/data-scientist.md +38 -0
- package/src/commands/deep-learning-engineer.md +47 -0
- package/src/commands/end.md +49 -0
- package/src/commands/knowledge.md +24 -0
- package/src/commands/ml-engineer.md +42 -0
- package/src/commands/mlops-engineer.md +47 -0
- package/src/commands/notebook-walkthrough.md +58 -0
- package/src/commands/researcher.md +40 -0
- package/src/commands/resume.md +57 -0
- package/src/commands/review-pr.md +26 -0
- package/src/commands/shards-guide.md +41 -0
- package/src/commands/shards-ui.md +32 -0
- package/src/commands/shards.md +41 -0
- package/src/docs/01-getting-started/concepts.md +109 -0
- package/src/docs/01-getting-started/first-session.md +79 -0
- package/src/docs/01-getting-started/install.md +61 -0
- package/src/docs/02-agents/academic.md +71 -0
- package/src/docs/02-agents/ai-engineer.md +78 -0
- package/src/docs/02-agents/analytics-engineer.md +58 -0
- package/src/docs/02-agents/applied-ml-scientist.md +59 -0
- package/src/docs/02-agents/backend-engineer.md +58 -0
- package/src/docs/02-agents/bi-engineer.md +65 -0
- package/src/docs/02-agents/data-analyst.md +67 -0
- package/src/docs/02-agents/data-engineer.md +57 -0
- package/src/docs/02-agents/data-modeller.md +51 -0
- package/src/docs/02-agents/data-scientist.md +78 -0
- package/src/docs/02-agents/deep-learning-engineer.md +64 -0
- package/src/docs/02-agents/ml-engineer.md +80 -0
- package/src/docs/02-agents/mlops-engineer.md +59 -0
- package/src/docs/02-agents/overview.md +62 -0
- package/src/docs/02-agents/researcher.md +73 -0
- package/src/docs/02-agents/syn.md +88 -0
- package/src/docs/03-protocols/auto-verify.md +82 -0
- package/src/docs/03-protocols/autonomous-research.md +59 -0
- package/src/docs/03-protocols/behavioral-rules.md +35 -0
- package/src/docs/03-protocols/diverge.md +50 -0
- package/src/docs/03-protocols/engineering-guidelines.md +56 -0
- package/src/docs/03-protocols/experiment-versioning.md +38 -0
- package/src/docs/03-protocols/gate-pattern.md +65 -0
- package/src/docs/03-protocols/incremental-testing.md +68 -0
- package/src/docs/03-protocols/join-path.md +46 -0
- package/src/docs/03-protocols/knowledge-ledger.md +70 -0
- package/src/docs/03-protocols/reviewer-verdicts.md +39 -0
- package/src/docs/03-protocols/swarm.md +40 -0
- package/src/docs/03-protocols/validation.md +174 -0
- package/src/docs/04-ui/activity-bar.md +70 -0
- package/src/docs/04-ui/chat-pane.md +80 -0
- package/src/docs/04-ui/code-intel.md +62 -0
- package/src/docs/04-ui/file-editing.md +61 -0
- package/src/docs/04-ui/git.md +54 -0
- package/src/docs/04-ui/keybindings.md +79 -0
- package/src/docs/04-ui/knowledge-map.md +76 -0
- package/src/docs/04-ui/overview.md +93 -0
- package/src/docs/04-ui/panels.md +49 -0
- package/src/docs/04-ui/pinboard-selection.md +66 -0
- package/src/docs/04-ui/quick-open-palette.md +56 -0
- package/src/docs/04-ui/sessions.md +81 -0
- package/src/docs/04-ui/settings-permissions.md +56 -0
- package/src/docs/05-commands/reference.md +59 -0
- package/src/docs/06-outputs/directory-map.md +116 -0
- package/src/docs/07-workflows/ai-eval-first.md +57 -0
- package/src/docs/07-workflows/deep-study-to-production.md +76 -0
- package/src/docs/07-workflows/diverge-exploration.md +77 -0
- package/src/docs/07-workflows/quick-analysis.md +45 -0
- package/src/docs/08-integrations/claude-code-auto-mode.md +191 -0
- package/src/docs/08-integrations/google-slides.md +175 -0
- package/src/docs/README.md +30 -0
- package/src/docs/manifest.json +108 -0
- package/src/templates/analysis-template.md +20 -0
- package/src/templates/branch-report.md +46 -0
- package/src/templates/diff-report.md +88 -0
- package/src/templates/knowledge-index.md +7 -0
- package/src/templates/model-card-schema.json +186 -0
- package/src/templates/model-card-schema.md +88 -0
- package/src/templates/model-card.md +124 -0
- package/src/templates/project-plan.md +47 -0
- package/src/templates/project-specs.md +81 -0
- package/src/templates/report-template.md +43 -0
- package/src/templates/study-template.md +25 -0
- package/src/ui/cc-readonly.js +181 -0
- package/src/ui/chat-session.js +466 -0
- package/src/ui/css/base.css +136 -0
- package/src/ui/css/brainstorm.css +525 -0
- package/src/ui/css/chat.css +1405 -0
- package/src/ui/css/editor.css +546 -0
- package/src/ui/css/eval-dashboard.css +157 -0
- package/src/ui/css/experiment.css +237 -0
- package/src/ui/css/guide.css +186 -0
- package/src/ui/css/knowledge-map.css +383 -0
- package/src/ui/css/layout.css +431 -0
- package/src/ui/css/model-card.css +161 -0
- package/src/ui/css/notebook-walkthrough.css +271 -0
- package/src/ui/css/pr-review.css +403 -0
- package/src/ui/css/prompt-lab.css +325 -0
- package/src/ui/css/sessions.css +258 -0
- package/src/ui/css/sidebar.css +661 -0
- package/src/ui/css/terminal.css +113 -0
- package/src/ui/css/theme-light.css +542 -0
- package/src/ui/index.html +389 -0
- package/src/ui/js/agents.js +32 -0
- package/src/ui/js/bookmarks.js +230 -0
- package/src/ui/js/chat.js +1776 -0
- package/src/ui/js/code-intel.js +328 -0
- package/src/ui/js/command-palette.js +142 -0
- package/src/ui/js/events.js +591 -0
- package/src/ui/js/explorer.js +317 -0
- package/src/ui/js/file-view.js +477 -0
- package/src/ui/js/git.js +536 -0
- package/src/ui/js/guide.js +198 -0
- package/src/ui/js/hud.js +75 -0
- package/src/ui/js/init.js +351 -0
- package/src/ui/js/knowledge-map.js +906 -0
- package/src/ui/js/markdown.js +114 -0
- package/src/ui/js/monaco.js +164 -0
- package/src/ui/js/notebook-walkthrough.js +272 -0
- package/src/ui/js/notebook.js +448 -0
- package/src/ui/js/panels.js +2681 -0
- package/src/ui/js/pinboard.js +186 -0
- package/src/ui/js/quick-open.js +164 -0
- package/src/ui/js/selection-context.js +131 -0
- package/src/ui/js/sessions.js +256 -0
- package/src/ui/js/settings.js +476 -0
- package/src/ui/js/split-view.js +82 -0
- package/src/ui/js/state.js +343 -0
- package/src/ui/js/table.js +161 -0
- package/src/ui/js/tabs.js +284 -0
- package/src/ui/js/tabular.js +125 -0
- package/src/ui/js/terminal.js +354 -0
- package/src/ui/js/timeline.js +137 -0
- package/src/ui/js/utils.js +293 -0
- package/src/ui/notebook-kernel.py +790 -0
- package/src/ui/open-browser.js +55 -0
- package/src/ui/permission-pattern.js +42 -0
- package/src/ui/relay.js +513 -0
- package/src/ui/server.js +3072 -0
- package/src/ui/session-index.js +225 -0
- package/src/ui/shards_icon.png +0 -0
- package/src/ui/spawn-server.js +41 -0
- package/src/ui/symbol-index.js +813 -0
- package/src/ui/ui-push.js +177 -0
- package/tools/gate-hook/VALIDATION_SPEC.md +273 -0
- package/tools/gate-hook/__tests__/auto-verify.test.js +343 -0
- package/tools/gate-hook/auto-allowlist.js +179 -0
- package/tools/gate-hook/auto-state.js +68 -0
- package/tools/gate-hook/classify.js +21 -0
- package/tools/gate-hook/log.js +57 -0
- package/tools/gate-hook/parser.js +205 -0
- package/tools/gate-hook/sql-guard.js +230 -0
- package/tools/gate-hook/state.js +170 -0
- package/tools/gate-hook/sweep.js +139 -0
- package/tools/gate-hook/transcript.js +45 -0
- package/tools/gate-hook/validation.js +321 -0
- package/tools/gate-hook.js +475 -0
- package/tools/install.js +914 -0
- package/tools/shards-gates.js +311 -0
- package/tools/shards-sessions.js +261 -0
- package/tools/shards-ui.js +377 -0
|
@@ -0,0 +1,170 @@
|
|
|
1
|
+
> **Previous:** phase-6.md confirmed
|
|
2
|
+
> **Next:** phase-7.md (read only after this phase's gate is confirmed)
|
|
3
|
+
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## Phase 6.5 — Winner Selection (Optional)
|
|
7
|
+
|
|
8
|
+
Goal: Finalize `eval-results.json` `bestCandidate` before Phase 7 generates the
|
|
9
|
+
model card and report. Phase 6 produces baseline + candidates; Phase 6.5 decides
|
|
10
|
+
who wins. The available selection paths depend on what Phase 6 actually produced.
|
|
11
|
+
|
|
12
|
+
This phase is **optional**. Skipping it leaves `bestCandidate` as whatever Phase
|
|
13
|
+
6 wrote (or unset), and Phase 7 proceeds accordingly.
|
|
14
|
+
|
|
15
|
+
### State inspection
|
|
16
|
+
|
|
17
|
+
Before presenting options, inspect:
|
|
18
|
+
- `<project_dir>/eval-results.json` — `status`, `dimensions[*].actual`, `bestCandidate`
|
|
19
|
+
- Candidate count documented in Phase 4
|
|
20
|
+
- Infrastructure dimensions (model size, inference time, memory) — pass/fail
|
|
21
|
+
|
|
22
|
+
Classify Phase 6 state into one of:
|
|
23
|
+
|
|
24
|
+
| State | Indicator |
|
|
25
|
+
|-------|-----------|
|
|
26
|
+
| **Not executed** | `status == "running"`, actuals null, notebook written but unrun |
|
|
27
|
+
| **Single evaluated** | 1 candidate with metrics filled, infrastructure pass |
|
|
28
|
+
| **Clear winner** | 2+ candidates evaluated, one dominates primary metric within budgets |
|
|
29
|
+
| **Mixed trade-offs** | 2+ candidates evaluated, no single dominant (e.g., one wins accuracy, another wins latency) |
|
|
30
|
+
|
|
31
|
+
### Present applicable options
|
|
32
|
+
|
|
33
|
+
Open with a state summary, then offer only the paths that apply to this state.
|
|
34
|
+
Do not present options the state does not support.
|
|
35
|
+
|
|
36
|
+
**Path availability matrix:**
|
|
37
|
+
|
|
38
|
+
| State | [D] Deterministic | [A] Solo AR | [F] AR fan-out | [M] Run manually | [S] Skip |
|
|
39
|
+
|-------|:-:|:-:|:-:|:-:|:-:|
|
|
40
|
+
| Not executed | — | ✓ | ✓ | ✓ | ✓ (with warning) |
|
|
41
|
+
| Single evaluated | ✓ | ✓ | — | — | ✓ |
|
|
42
|
+
| Clear winner | ✓ | ✓ | ✓ | — | ✓ |
|
|
43
|
+
| Mixed trade-offs | ✓ (user-prioritized metric) | ✓ | ✓ | — | ✓ |
|
|
44
|
+
|
|
45
|
+
Note: `[F]` AR fan-out requires 2+ mutually exclusive candidate families from
|
|
46
|
+
Phase 4. If only one family exists, fan-out is unavailable.
|
|
47
|
+
|
|
48
|
+
**Presentation format:**
|
|
49
|
+
|
|
50
|
+
```
|
|
51
|
+
**Phase 6.5 — Winner Selection**
|
|
52
|
+
|
|
53
|
+
Phase 6 state: <classification>
|
|
54
|
+
<one-line summary: candidate count, metrics available, infrastructure pass/fail>
|
|
55
|
+
|
|
56
|
+
Available paths:
|
|
57
|
+
[D] Deterministic pick — lock the leader on primary metric. No iteration work.
|
|
58
|
+
Cost: none. Available when candidates have been evaluated.
|
|
59
|
+
[A] Solo AR on leading candidate — push the leader further for N iterations.
|
|
60
|
+
Cost: N iterations + token/compute spend. Requires git/DVC.
|
|
61
|
+
[F] AR fan-out across approach families — K parallel AR loops, arbiter picks.
|
|
62
|
+
Cost: K × N iterations + token/compute spend. Requires git/DVC and cost ceiling.
|
|
63
|
+
[M] I'll run the notebook myself — come back with results.
|
|
64
|
+
Cost: none to agent. Re-enters 6.5 after you return.
|
|
65
|
+
[S] Skip winner selection — lock whatever is in eval-results.json and move on.
|
|
66
|
+
If bestCandidate is null, you'll need to name the winner explicitly.
|
|
67
|
+
|
|
68
|
+
Which path?
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
Wait for the user's selection.
|
|
72
|
+
|
|
73
|
+
### Execute the selected path
|
|
74
|
+
|
|
75
|
+
**[D] Deterministic pick:**
|
|
76
|
+
1. Read `eval-results.json`.
|
|
77
|
+
2. Select the candidate with best value on the primary metric (max or min per
|
|
78
|
+
Phase 4 direction).
|
|
79
|
+
3. Verify infrastructure dimensions (model size, inference time, memory) PASS
|
|
80
|
+
for this candidate.
|
|
81
|
+
4. If infrastructure FAILs, surface: "The metric winner fails <dimension>.
|
|
82
|
+
Options: (a) accept and document the trade-off, (b) pick next-best candidate
|
|
83
|
+
that passes, (c) stop and revisit Phase 5 budgets." Wait for decision.
|
|
84
|
+
5. Update `eval-results.json`: set `bestCandidate.model`, `bestCandidate.metrics`,
|
|
85
|
+
`bestCandidate.deltas` (vs. baseline). Recompute `summary.overallVerdict`.
|
|
86
|
+
|
|
87
|
+
**[A] Solo AR on leading candidate:**
|
|
88
|
+
1. Pre-flight: verify git (or DVC) is available per
|
|
89
|
+
`.claude/agents/specific_instructions/shared/experiment_versioning.md`
|
|
90
|
+
Section A. If unavailable, refuse and offer `[D]` instead.
|
|
91
|
+
2. Determine the leading candidate (same selection logic as `[D]`). If no
|
|
92
|
+
candidate has been evaluated, use the Phase 4 candidate the user identifies.
|
|
93
|
+
3. Read `.claude/agents/specific_instructions/ml_engineer/research.md` and
|
|
94
|
+
execute Sections A-G with abbreviated setup:
|
|
95
|
+
- Primary metric, direction, target: inherit from Phase 4
|
|
96
|
+
- Baseline value: the leading candidate's current metric (from Phase 6)
|
|
97
|
+
- Mutable scope: inherit from Phase 5 training pipeline
|
|
98
|
+
- Immutable scope: `data/`, `eval/`, `deploy/` by default
|
|
99
|
+
- Preset: ask user (interactive | overnight)
|
|
100
|
+
4. Announce the AR behavioral exception and run the loop.
|
|
101
|
+
5. On convergence: write converged metrics into `eval-results.json`
|
|
102
|
+
`bestCandidate`. Record iterations consumed and budget spent in Phase 6.5 doc.
|
|
103
|
+
|
|
104
|
+
**[F] AR fan-out across approach families:**
|
|
105
|
+
1. Pre-flight: verify git/DVC available and cost ceiling confirmed.
|
|
106
|
+
2. Identify 2-3 approach families from Phase 4 candidates. Confirm mutual
|
|
107
|
+
exclusivity and genuine viability with the user.
|
|
108
|
+
3. Follow
|
|
109
|
+
`.claude/agents/specific_instructions/shared/diverge_protocol.md` Section B
|
|
110
|
+
(DIVERGE proposal gate — use ID
|
|
111
|
+
`specific-instructions-shared-diverge-protocol-ar-<project>`).
|
|
112
|
+
4. On confirmation, follow `autonomous_research.md` Section H:
|
|
113
|
+
- H.3 for parallel branch spawning (all branches in one message)
|
|
114
|
+
- H.5 for git strategy (default `branch-local`)
|
|
115
|
+
- Wait for all branches to complete
|
|
116
|
+
5. Run arbiter per Section F of `diverge_protocol.md`.
|
|
117
|
+
6. On user winner selection: promote per Section G — copy artifacts, squash-merge
|
|
118
|
+
on chosen git strategy, update `eval-results.json.bestCandidate` with the
|
|
119
|
+
winning branch's converged metrics.
|
|
120
|
+
|
|
121
|
+
**[M] User runs notebook manually:**
|
|
122
|
+
1. Tell the user: "Run <notebook path>. Populate the metrics tables, then tell
|
|
123
|
+
me when you're back. I'll re-inspect state and re-present options."
|
|
124
|
+
2. Wait for user return. Do not advance.
|
|
125
|
+
3. On return: re-run state inspection and re-enter the option presentation.
|
|
126
|
+
|
|
127
|
+
**[S] Skip:**
|
|
128
|
+
1. If `bestCandidate` is populated: "Proceeding to Phase 7 with
|
|
129
|
+
`bestCandidate = <model type>`, primary metric = <value>. Infrastructure:
|
|
130
|
+
<pass/fail>. Confirm?"
|
|
131
|
+
2. If `bestCandidate` is null: "No winner is recorded. Which candidate should I
|
|
132
|
+
write as the winner before Phase 7?" Wait for the user to name one, then
|
|
133
|
+
populate `bestCandidate` from the corresponding Phase 6 entry.
|
|
134
|
+
|
|
135
|
+
### Document Phase 6.5
|
|
136
|
+
|
|
137
|
+
Append to `project-specs.md`:
|
|
138
|
+
|
|
139
|
+
```markdown
|
|
140
|
+
---
|
|
141
|
+
|
|
142
|
+
## Phase 6.5: Winner Selection (ML Engineer)
|
|
143
|
+
- **Phase 6 state at entry:** Not executed | Single evaluated | Clear winner | Mixed trade-offs
|
|
144
|
+
- **Path selected:** Deterministic | Solo AR | AR fan-out | Manual | Skip
|
|
145
|
+
- **Rationale:** <one-line reason — e.g., "clear metric winner, no need to push further">
|
|
146
|
+
- **Winner:**
|
|
147
|
+
- Model type: <type>
|
|
148
|
+
- Primary metric: <metric> = <value>
|
|
149
|
+
- Infrastructure: model size <X>MB, inference <X>ms, memory <X>MB — Pass | Fail
|
|
150
|
+
- Deltas vs. baseline: <metric>: <delta>
|
|
151
|
+
- **AR details (if Solo AR or fan-out):**
|
|
152
|
+
- Preset: interactive | overnight | custom
|
|
153
|
+
- Iterations run: <N>
|
|
154
|
+
- Budget consumed: <tokens | dollars>
|
|
155
|
+
- Convergence: converged | budget exhausted | user interrupted
|
|
156
|
+
- Branches (fan-out only):
|
|
157
|
+
- `<branch-slug>`: <converged metric value> — <winner | runner-up>
|
|
158
|
+
- **Infrastructure trade-off (if applicable):** <metric winner fails <dim> — user accepted | chose next-best | revisited Phase 5>
|
|
159
|
+
- **eval-results.json bestCandidate:** Locked | Skipped — bestCandidate left as-is
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
::GATE:: id=ml-engineer-phase-6-5 phase=6.5 kind=phase
|
|
163
|
+
Read this section back to the user. Stop here — do not begin Phase 7 or output any further content. Wait for the user to explicitly confirm before proceeding.
|
|
164
|
+
::ENDGATE::
|
|
165
|
+
|
|
166
|
+
---
|
|
167
|
+
|
|
168
|
+
## When this gate is confirmed
|
|
169
|
+
|
|
170
|
+
Read `.claude/agents/specific_instructions/ml_engineer/phases/phase-7.md` in full and follow its instructions starting from Phase 7.
|
|
@@ -0,0 +1,295 @@
|
|
|
1
|
+
> **Previous:** phase-5.md confirmed
|
|
2
|
+
> **Next:** phase-6-5.md if Phase 6.5 is applicable per the state inspection; otherwise phase-7.md (read only after this phase's gate is confirmed)
|
|
3
|
+
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## Phase 6 — Execute
|
|
7
|
+
|
|
8
|
+
**Context checkpoint:** Before building, prompt the user:
|
|
9
|
+
|
|
10
|
+
"Planning's locked — good moment to run `/compact` or `/clear` before we start
|
|
11
|
+
executing. I'll be working from project-specs.md from here. Say the word when
|
|
12
|
+
you're ready."
|
|
13
|
+
|
|
14
|
+
Wait for any signal from the user before beginning build steps.
|
|
15
|
+
|
|
16
|
+
**Knowledge re-check:** Follow `.claude/agents/specific_instructions/shared/knowledge_checkpoint.md` before building.
|
|
17
|
+
|
|
18
|
+
Goal: Build the feature queries, training notebook, and pipeline artifacts.
|
|
19
|
+
|
|
20
|
+
**Join path self-check (feature queries):** Before requesting the Data Modeller
|
|
21
|
+
review, trace the join path for each feature query following
|
|
22
|
+
`.claude/agents/specific_instructions/shared/join_path_protocol.md`. Present the
|
|
23
|
+
trace to the user. Include it in the DM prompt below.
|
|
24
|
+
|
|
25
|
+
**Then request Data Modeller query review with validation:**
|
|
26
|
+
|
|
27
|
+
Tell the user: "Pulling in the Data Modeller to verify the feature extraction queries. Feature pipeline built on bad grain assumptions is a training set problem. I'm not building until this is confirmed."
|
|
28
|
+
|
|
29
|
+
```
|
|
30
|
+
Task(
|
|
31
|
+
subagent_type="data-modeller",
|
|
32
|
+
description="Review ML feature queries for [project]",
|
|
33
|
+
prompt="I am the ML Engineer shard. I've written feature extraction queries for
|
|
34
|
+
project [name]. The project specs are at: [services|<existing_dir>]/[name]/project-specs.md
|
|
35
|
+
|
|
36
|
+
Here are the queries:
|
|
37
|
+
[include query outlines or key SQL]
|
|
38
|
+
|
|
39
|
+
Please REVIEW (not just explore): Do the joins make sense given the data model
|
|
40
|
+
grain? Are there grain fan-out risks? Am I using the right tables for these
|
|
41
|
+
features?
|
|
42
|
+
|
|
43
|
+
Run validation queries to check:
|
|
44
|
+
1. PK uniqueness on all tables referenced in these queries
|
|
45
|
+
2. Null rates on join keys and key feature source columns
|
|
46
|
+
3. Join fan-out: row counts before/after the joins in my feature queries
|
|
47
|
+
4. Data freshness on the tables feeding features
|
|
48
|
+
|
|
49
|
+
Cross-reference against the project requirements in project-specs.md
|
|
50
|
+
(especially Phase 3 feature candidates and Phase 5 data freshness requirements).
|
|
51
|
+
This is for ML feature engineering — pay special attention to fan-out that would
|
|
52
|
+
silently inflate training examples.
|
|
53
|
+
Return your full review with query validation results."
|
|
54
|
+
)
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
The Data Modeller will run a bulk read-only validation sweep — that's an
|
|
58
|
+
auto-verify fit on the consultation side (the Data Modeller's `service_mode.md`
|
|
59
|
+
already references it). On your side, no marker is needed for the Task call
|
|
60
|
+
itself; auto-verify only matters when the calling agent is the one running
|
|
61
|
+
the bulk queries.
|
|
62
|
+
|
|
63
|
+
**Then build:**
|
|
64
|
+
|
|
65
|
+
1. **SQL queries** — Write to:
|
|
66
|
+
- Greenfield: `models/<name>/queries/`
|
|
67
|
+
- Iteration: `<existing_service_dir>/queries/`
|
|
68
|
+
- Name files descriptively: `01_label_definition.sql`, `02_user_features.sql`,
|
|
69
|
+
`03_behavioral_features.sql`, `04_training_dataset.sql`
|
|
70
|
+
- Include header comments:
|
|
71
|
+
```sql
|
|
72
|
+
-- Project: <project_name>
|
|
73
|
+
-- Query: <description>
|
|
74
|
+
-- Date: <date>
|
|
75
|
+
-- Feature group: <label | user | behavioral | contextual | interaction>
|
|
76
|
+
-- Dependencies: <upstream tables>
|
|
77
|
+
-- Output grain: one row per <entity>
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
2. **Training notebook** — Write using NotebookEdit to:
|
|
81
|
+
- Greenfield: `models/<name>/notebooks/`
|
|
82
|
+
- Iteration: `<existing_service_dir>/notebooks/`
|
|
83
|
+
Structure:
|
|
84
|
+
- **SQL loading rule** — **Do NOT re-embed SQL as Python strings.** Read `.sql`
|
|
85
|
+
files directly using `Path.read_text()`. Reference files by relative path from
|
|
86
|
+
the notebook location:
|
|
87
|
+
```python
|
|
88
|
+
from pathlib import Path
|
|
89
|
+
sql = Path("../queries/02_user_features.sql").read_text()
|
|
90
|
+
df = pd.read_sql(sql, conn)
|
|
91
|
+
```
|
|
92
|
+
- **Overview** (markdown): business problem, model type, key decisions
|
|
93
|
+
- **Setup**: imports, config, random seeds, data loading
|
|
94
|
+
- **Feature Engineering**: feature computation, transformations, encoding
|
|
95
|
+
- **EDA**: target distribution, feature distributions, correlations, class balance
|
|
96
|
+
- **Baseline Model**: train, evaluate, establish floor
|
|
97
|
+
- **Candidate Model(s)**: train, tune, evaluate, compare to baseline
|
|
98
|
+
- **Model Analysis**: feature importance, SHAP values, error analysis
|
|
99
|
+
- **Infrastructure Readiness**: model size, inference time benchmarks,
|
|
100
|
+
serving requirements check
|
|
101
|
+
- **Results Summary**: final metrics, business interpretation, recommendation
|
|
102
|
+
|
|
103
|
+
3. **Requirements file** — `requirements.txt` with all ML dependencies
|
|
104
|
+
|
|
105
|
+
4. **Config file** (if applicable) — model hyperparameters, feature lists, thresholds
|
|
106
|
+
|
|
107
|
+
5. **Eval results JSON** — After training and evaluating models, write structured
|
|
108
|
+
results to the project's `eval-results.json`:
|
|
109
|
+
- Greenfield: `models/<name>/eval-results.json`
|
|
110
|
+
- Iteration: `<existing_service_dir>/eval-results.json`
|
|
111
|
+
|
|
112
|
+
The JSON must follow this schema:
|
|
113
|
+
```json
|
|
114
|
+
{
|
|
115
|
+
"variant": "ml-engineer",
|
|
116
|
+
"projectName": "<project_name>",
|
|
117
|
+
"status": "running",
|
|
118
|
+
"timestamp": "<ISO-8601>",
|
|
119
|
+
"summary": {
|
|
120
|
+
"totalDimensions": 0,
|
|
121
|
+
"passed": 0,
|
|
122
|
+
"failed": 0,
|
|
123
|
+
"overallVerdict": "PENDING"
|
|
124
|
+
},
|
|
125
|
+
"dimensions": [
|
|
126
|
+
{ "dimension": "<metric_name>", "metric": "<metric>", "target": 0.85, "actual": null, "unit": "ratio", "verdict": null }
|
|
127
|
+
],
|
|
128
|
+
"cost": {
|
|
129
|
+
"perRequest": null,
|
|
130
|
+
"per1kTokens": null,
|
|
131
|
+
"monthlyProjected": null,
|
|
132
|
+
"budget": null,
|
|
133
|
+
"currency": "USD"
|
|
134
|
+
},
|
|
135
|
+
"baseline": {
|
|
136
|
+
"model": "<model_type>",
|
|
137
|
+
"metrics": { "<metric>": null }
|
|
138
|
+
},
|
|
139
|
+
"bestCandidate": {
|
|
140
|
+
"model": "<model_type>",
|
|
141
|
+
"metrics": { "<metric>": null },
|
|
142
|
+
"deltas": { "<metric>": null }
|
|
143
|
+
},
|
|
144
|
+
"infrastructure": [
|
|
145
|
+
{ "dimension": "Model size", "actual": null, "budget": "<Y>MB", "verdict": null },
|
|
146
|
+
{ "dimension": "Inference time", "actual": null, "budget": "<Y>ms", "verdict": null },
|
|
147
|
+
{ "dimension": "Memory usage", "actual": null, "budget": "<Y>MB", "verdict": null }
|
|
148
|
+
]
|
|
149
|
+
}
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
Write the file initially with `status: "running"` and null values. Update it
|
|
153
|
+
as baseline and candidate models are evaluated. When complete, set
|
|
154
|
+
`status: "complete"`, populate all metrics, compute `summary.overallVerdict`
|
|
155
|
+
(PASS if all dimensions pass, FAIL if any fail, PARTIAL if mixed).
|
|
156
|
+
|
|
157
|
+
If the Shards UI is active (`.shards/ui.port` file exists), push the eval
|
|
158
|
+
dashboard panel:
|
|
159
|
+
```bash
|
|
160
|
+
node .shards/ui/ui-push.js eval-dashboard \
|
|
161
|
+
--title "Eval: <project_name>" \
|
|
162
|
+
--agent "ml-engineer" \
|
|
163
|
+
--panel-id "eval-<project_name>" \
|
|
164
|
+
--source "<path_to>/eval-results.json"
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
6. **Applied ML Scientist results review** — After baseline and candidate models
|
|
168
|
+
are trained and evaluated, consult the Applied ML Scientist to review the
|
|
169
|
+
*actual results*, not just the design. AMS reviewed the plan at Phase 4; this
|
|
170
|
+
is where the plan meets empirical reality. Skip if Phase 4's AMS consultation
|
|
171
|
+
was skipped per the standard-tabular criteria.
|
|
172
|
+
|
|
173
|
+
Tell the user: "Getting the Applied ML Scientist back in to review the
|
|
174
|
+
training results — loss curves, eval metrics, model behavior. Design review
|
|
175
|
+
is cheap; results review catches what theory missed."
|
|
176
|
+
|
|
177
|
+
```
|
|
178
|
+
Task(
|
|
179
|
+
subagent_type="applied-ml-scientist",
|
|
180
|
+
description="Training results review for <project_name>",
|
|
181
|
+
prompt="I am the ML Engineer shard. I've trained baseline and candidate
|
|
182
|
+
models for <project_name>. You reviewed the design at Phase 4 — now I need
|
|
183
|
+
you to review the actual results.
|
|
184
|
+
|
|
185
|
+
Project specs: <path to project-specs.md>
|
|
186
|
+
Eval results: <path to eval-results.json>
|
|
187
|
+
Notebook: <path to training notebook>
|
|
188
|
+
|
|
189
|
+
Baseline results:
|
|
190
|
+
- Model: <type>
|
|
191
|
+
- Key metrics: <metric>: <value>, <metric>: <value>
|
|
192
|
+
|
|
193
|
+
Best candidate results:
|
|
194
|
+
- Model: <type>
|
|
195
|
+
- Key metrics: <metric>: <value>, <metric>: <value>
|
|
196
|
+
- Delta vs baseline: <delta>
|
|
197
|
+
|
|
198
|
+
Training artifacts available:
|
|
199
|
+
- Loss curves: <notebook cell or plot path>
|
|
200
|
+
- Learning curves (train vs val): <notebook cell or plot path>
|
|
201
|
+
- Error analysis: <notebook cell or plot path if present>
|
|
202
|
+
- Feature importance / SHAP: <notebook cell or plot path if present>
|
|
203
|
+
|
|
204
|
+
Please review the actual results and flag:
|
|
205
|
+
1. Training dynamics — do loss/learning curves look healthy, or do you see
|
|
206
|
+
signs of underfitting, overfitting, instability, or premature convergence?
|
|
207
|
+
2. Result plausibility — does the baseline-to-candidate lift match what
|
|
208
|
+
the methodology predicted, or is something suspicious (too-good-to-be-true
|
|
209
|
+
metrics, unexplained gaps between validation and test)?
|
|
210
|
+
3. Error structure — does the error analysis reveal any systematic failure
|
|
211
|
+
modes that suggest a methodology adjustment?
|
|
212
|
+
4. Are there concrete next iterations from the literature that would address
|
|
213
|
+
observable gaps in current performance?
|
|
214
|
+
|
|
215
|
+
Return a structured ML Science Review per your service-mode format."
|
|
216
|
+
)
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
Apply the Reviewer Verdict Protocol. Append the review to project-specs.md
|
|
220
|
+
under Phase 6.
|
|
221
|
+
|
|
222
|
+
### Incremental testing — checkpoint gates between components
|
|
223
|
+
|
|
224
|
+
Follow `.claude/agents/specific_instructions/shared/incremental_testing.md` during this build. Each component above is a checkpoint seam — after you write and execute a component, emit a `kind=checkpoint` gate fence (template below) and wait for user confirmation before starting the next component. Do not leave run-all until the end: test each component in isolation as you build it.
|
|
225
|
+
|
|
226
|
+
Checkpoint gate fence — emit exactly this shape. Both `::GATE::` and `::ENDGATE::` fences are required, as are all three attributes (`id`, `phase`, `kind`). No prose outside the fence.
|
|
227
|
+
|
|
228
|
+
```
|
|
229
|
+
::GATE:: id=<agent-name>-phase-<N>-checkpoint-<component> phase=<N> kind=checkpoint
|
|
230
|
+
Component: <human-readable name>
|
|
231
|
+
Test command: <exact command you ran>
|
|
232
|
+
Evidence:
|
|
233
|
+
- <measured fact 1, e.g. "df.shape = (48211, 47)">
|
|
234
|
+
- <measured fact 2, e.g. "null rate on join key = 0.00%">
|
|
235
|
+
- <measured fact 3, e.g. "sample head matches expected schema">
|
|
236
|
+
Status: PASS | FAIL — <one-line summary>
|
|
237
|
+
Next: <what you'll build after this is confirmed>
|
|
238
|
+
Stop here — await explicit confirmation before writing the next component.
|
|
239
|
+
::ENDGATE::
|
|
240
|
+
```
|
|
241
|
+
|
|
242
|
+
Expected checkpoint gate IDs for this phase (emit in order as you build):
|
|
243
|
+
|
|
244
|
+
- `ml-engineer-phase-6-checkpoint-queries` — each feature query returns expected shape under `LIMIT 100`; row counts before/after joins match Phase 3 prediction.
|
|
245
|
+
- `ml-engineer-phase-6-checkpoint-data` — notebook data-load + EDA cells produce expected shape; target distribution matches prior knowledge.
|
|
246
|
+
- `ml-engineer-phase-6-checkpoint-baseline` — baseline model fits and evaluates on held-out data; metrics recorded.
|
|
247
|
+
- `ml-engineer-phase-6-checkpoint-candidate` — candidate model smoke-fits on ≤1% of data first (loss decreasing), then full fit; eval metrics beat baseline floor.
|
|
248
|
+
- `ml-engineer-phase-6-checkpoint-infra` — model size, inference latency, memory measured against Phase 5 budget.
|
|
249
|
+
|
|
250
|
+
The hook blocks all non-read tools while a checkpoint is open. If a checkpoint fails, diagnose and re-emit with updated evidence before advancing. Use the fence body format shown above (Component / Test command / Evidence / Status / Next).
|
|
251
|
+
|
|
252
|
+
### Document Phase 6
|
|
253
|
+
|
|
254
|
+
```markdown
|
|
255
|
+
---
|
|
256
|
+
|
|
257
|
+
## Phase 6: Build Log (ML Engineer)
|
|
258
|
+
- **Data Modeller query review:**
|
|
259
|
+
- Verdict: Approved | Concerns raised
|
|
260
|
+
- Notes: <summary>
|
|
261
|
+
- Issues addressed: <how resolved or "none raised">
|
|
262
|
+
- **Applied ML Scientist results review:** <summary if consulted> | Skipped — Phase 4 AMS review was skipped per standard-tabular criteria
|
|
263
|
+
- Verdict: Sound | Consider Alternatives | Revise
|
|
264
|
+
- Tier: Proceed | Proceed with caveats | Halt
|
|
265
|
+
- Training dynamics notes: <healthy | concerns — description>
|
|
266
|
+
- Error structure notes: <systematic failures surfaced or "none">
|
|
267
|
+
- Next-iteration suggestions: <from AMS or "none">
|
|
268
|
+
- Reviewer resolution: Approved | Approved on resubmit | User override — <rationale> | Project stopped
|
|
269
|
+
- **Query files:**
|
|
270
|
+
- <file path>: <description>
|
|
271
|
+
- **Notebook location:** <file path>
|
|
272
|
+
- **Requirements file:** <file path>
|
|
273
|
+
- **Config file:** <file path or "N/A">
|
|
274
|
+
- **Baseline results:**
|
|
275
|
+
- <metric>: <value>
|
|
276
|
+
- **Best candidate results:**
|
|
277
|
+
- Model: <type>
|
|
278
|
+
- <metric>: <value> (improvement over baseline: <delta>)
|
|
279
|
+
- **Infrastructure readiness:**
|
|
280
|
+
- Model size: <X>MB (budget: <Y>MB) — Pass | Fail
|
|
281
|
+
- Inference time: <X>ms (budget: <Y>ms) — Pass | Fail
|
|
282
|
+
- Memory usage: <X>MB (budget: <Y>MB) — Pass | Fail
|
|
283
|
+
- **Deviations from plan:** <changes and why, or "none">
|
|
284
|
+
- **Surprising findings:** <anything unexpected>
|
|
285
|
+
```
|
|
286
|
+
|
|
287
|
+
::GATE:: id=ml-engineer-phase-6 phase=6 kind=phase validates=ml_engineer
|
|
288
|
+
Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
|
|
289
|
+
::ENDGATE::
|
|
290
|
+
|
|
291
|
+
---
|
|
292
|
+
|
|
293
|
+
## When this gate is confirmed
|
|
294
|
+
|
|
295
|
+
If Phase 6.5 is applicable for this project (based on the state classification in this phase), read `.claude/agents/specific_instructions/ml_engineer/phases/phase-6-5.md` in full and follow its instructions. Otherwise, read `.claude/agents/specific_instructions/ml_engineer/phases/phase-7.md` in full and follow its instructions starting from Phase 7. Do not pre-read further phase files.
|