@proflandrigan/shards 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +475 -0
- package/package.json +37 -0
- package/src/agents/academic.md +276 -0
- package/src/agents/ai-engineer.md +377 -0
- package/src/agents/analytics-engineer.md +364 -0
- package/src/agents/applied-ml-scientist.md +410 -0
- package/src/agents/backend-engineer.md +255 -0
- package/src/agents/bi-engineer.md +333 -0
- package/src/agents/data-analyst.md +343 -0
- package/src/agents/data-engineer.md +260 -0
- package/src/agents/data-modeller.md +386 -0
- package/src/agents/data-scientist.md +366 -0
- package/src/agents/deep-learning-engineer.md +389 -0
- package/src/agents/ml-engineer.md +424 -0
- package/src/agents/mlops-engineer.md +339 -0
- package/src/agents/researcher.md +187 -0
- package/src/agents/specific_instructions/academic/critical_review.md +263 -0
- package/src/agents/specific_instructions/academic/report.md +113 -0
- package/src/agents/specific_instructions/ai_engineer/advise.md +162 -0
- package/src/agents/specific_instructions/ai_engineer/bi_engineer_handoff.md +86 -0
- package/src/agents/specific_instructions/ai_engineer/experiment.md +471 -0
- package/src/agents/specific_instructions/ai_engineer/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/ai_engineer/phases/index.md +45 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-1.md +55 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-2.md +86 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-3.md +96 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-4.md +138 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-5.md +157 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-6.md +196 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-7.md +313 -0
- package/src/agents/specific_instructions/ai_engineer/phases.md +1011 -0
- package/src/agents/specific_instructions/ai_engineer/prompt_lab.md +161 -0
- package/src/agents/specific_instructions/ai_engineer/prompt_lab_ui_mode.md +28 -0
- package/src/agents/specific_instructions/ai_engineer/research.md +393 -0
- package/src/agents/specific_instructions/ai_engineer/research_ui_mode.md +66 -0
- package/src/agents/specific_instructions/ai_engineer/review.md +159 -0
- package/src/agents/specific_instructions/ai_engineer/validation_checklist.md +182 -0
- package/src/agents/specific_instructions/analytics_engineer/advise.md +155 -0
- package/src/agents/specific_instructions/analytics_engineer/bi_engineer_handoff.md +91 -0
- package/src/agents/specific_instructions/analytics_engineer/data_analyst_handoff.md +84 -0
- package/src/agents/specific_instructions/analytics_engineer/deep_phases.md +818 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/index.md +24 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-1.md +77 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-2.md +106 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-3.md +93 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-4.md +79 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-5.md +61 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-6.md +45 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-7.md +235 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-8.md +221 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-1.md +47 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-2.md +78 -0
- package/src/agents/specific_instructions/analytics_engineer/quick_phases.md +112 -0
- package/src/agents/specific_instructions/analytics_engineer/review.md +167 -0
- package/src/agents/specific_instructions/analytics_engineer/service_mode.md +369 -0
- package/src/agents/specific_instructions/analytics_engineer/ui_mode.md +45 -0
- package/src/agents/specific_instructions/analytics_engineer/update.md +162 -0
- package/src/agents/specific_instructions/analytics_engineer/validation_checklist.md +121 -0
- package/src/agents/specific_instructions/applied_ml_scientist/advise.md +143 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/index.md +21 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-1.md +51 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-2.md +66 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-3.md +113 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-4.md +104 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-5.md +156 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases.md +428 -0
- package/src/agents/specific_instructions/applied_ml_scientist/research.md +379 -0
- package/src/agents/specific_instructions/applied_ml_scientist/review.md +142 -0
- package/src/agents/specific_instructions/applied_ml_scientist/validation_checklist.md +136 -0
- package/src/agents/specific_instructions/backend_engineer/clean.md +149 -0
- package/src/agents/specific_instructions/backend_engineer/review.md +91 -0
- package/src/agents/specific_instructions/backend_engineer/review_checklist.md +54 -0
- package/src/agents/specific_instructions/backend_engineer/service_mode.md +67 -0
- package/src/agents/specific_instructions/bi_engineer/advise.md +137 -0
- package/src/agents/specific_instructions/bi_engineer/data_analyst_handoff.md +77 -0
- package/src/agents/specific_instructions/bi_engineer/incoming_handoff.md +45 -0
- package/src/agents/specific_instructions/bi_engineer/phases/index.md +20 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-1.md +164 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-2.md +92 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-3.md +121 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-4.md +106 -0
- package/src/agents/specific_instructions/bi_engineer/phases.md +451 -0
- package/src/agents/specific_instructions/bi_engineer/review.md +166 -0
- package/src/agents/specific_instructions/bi_engineer/update.md +147 -0
- package/src/agents/specific_instructions/bi_engineer/validation_checklist.md +124 -0
- package/src/agents/specific_instructions/data_analyst/advise.md +138 -0
- package/src/agents/specific_instructions/data_analyst/explain.md +221 -0
- package/src/agents/specific_instructions/data_analyst/incoming_handoff.md +40 -0
- package/src/agents/specific_instructions/data_analyst/phases/index.md +20 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-1.md +159 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-2.md +112 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-3.md +265 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-4.md +100 -0
- package/src/agents/specific_instructions/data_analyst/phases.md +501 -0
- package/src/agents/specific_instructions/data_analyst/review.md +138 -0
- package/src/agents/specific_instructions/data_analyst/ui_mode.md +26 -0
- package/src/agents/specific_instructions/data_analyst/update.md +144 -0
- package/src/agents/specific_instructions/data_analyst/validation_checklist.md +95 -0
- package/src/agents/specific_instructions/data_engineer/advise.md +137 -0
- package/src/agents/specific_instructions/data_engineer/phases.md +466 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/index.md +23 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-1.md +49 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-2.md +93 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-3.md +55 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-4.md +48 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-5.md +40 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-6.md +102 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-7.md +87 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/phase-1.md +45 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/phase-2.md +54 -0
- package/src/agents/specific_instructions/data_engineer/review.md +135 -0
- package/src/agents/specific_instructions/data_engineer/validation_checklist.md +136 -0
- package/src/agents/specific_instructions/data_modeller/advise.md +137 -0
- package/src/agents/specific_instructions/data_modeller/phases.md +581 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/index.md +23 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-1.md +52 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-2.md +113 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-3.md +47 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-4.md +51 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-5.md +45 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-6.md +105 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-7.md +136 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/phase-1.md +47 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/phase-2.md +65 -0
- package/src/agents/specific_instructions/data_modeller/review.md +141 -0
- package/src/agents/specific_instructions/data_modeller/service_mode.md +218 -0
- package/src/agents/specific_instructions/data_modeller/validation_checklist.md +125 -0
- package/src/agents/specific_instructions/data_scientist/advise.md +158 -0
- package/src/agents/specific_instructions/data_scientist/bi_engineer_handoff.md +63 -0
- package/src/agents/specific_instructions/data_scientist/experiment.md +482 -0
- package/src/agents/specific_instructions/data_scientist/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/data_scientist/explain.md +247 -0
- package/src/agents/specific_instructions/data_scientist/greenfield_data.md +35 -0
- package/src/agents/specific_instructions/data_scientist/ml_engineer_handoff.md +52 -0
- package/src/agents/specific_instructions/data_scientist/notebook_walkthrough.md +76 -0
- package/src/agents/specific_instructions/data_scientist/phases/index.md +24 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-1.md +45 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-2.md +67 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-3.md +89 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-4.md +143 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-5.md +71 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-6.md +239 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-7.md +207 -0
- package/src/agents/specific_instructions/data_scientist/phases.md +651 -0
- package/src/agents/specific_instructions/data_scientist/research.md +345 -0
- package/src/agents/specific_instructions/data_scientist/research_ui_mode.md +52 -0
- package/src/agents/specific_instructions/data_scientist/review.md +136 -0
- package/src/agents/specific_instructions/data_scientist/service_mode.md +247 -0
- package/src/agents/specific_instructions/data_scientist/validation_checklist.md +183 -0
- package/src/agents/specific_instructions/deep_learning_engineer/advise.md +145 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/index.md +21 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-1.md +74 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-2.md +98 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-3.md +76 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-4.md +128 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-5.md +292 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases.md +567 -0
- package/src/agents/specific_instructions/deep_learning_engineer/research.md +389 -0
- package/src/agents/specific_instructions/deep_learning_engineer/review.md +155 -0
- package/src/agents/specific_instructions/deep_learning_engineer/validation_checklist.md +147 -0
- package/src/agents/specific_instructions/ml_engineer/advise.md +174 -0
- package/src/agents/specific_instructions/ml_engineer/bi_engineer_handoff.md +71 -0
- package/src/agents/specific_instructions/ml_engineer/experiment.md +474 -0
- package/src/agents/specific_instructions/ml_engineer/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/ml_engineer/notebook_walkthrough.md +75 -0
- package/src/agents/specific_instructions/ml_engineer/phases/index.md +25 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-1.md +49 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-2.md +75 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-3.md +124 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-4.md +279 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-5.md +160 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-6-5.md +170 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-6.md +295 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-7.md +337 -0
- package/src/agents/specific_instructions/ml_engineer/phases.md +1068 -0
- package/src/agents/specific_instructions/ml_engineer/research.md +437 -0
- package/src/agents/specific_instructions/ml_engineer/research_ui_mode.md +71 -0
- package/src/agents/specific_instructions/ml_engineer/review.md +187 -0
- package/src/agents/specific_instructions/ml_engineer/service_mode.md +273 -0
- package/src/agents/specific_instructions/ml_engineer/validation_checklist.md +185 -0
- package/src/agents/specific_instructions/mlops_engineer/advise.md +139 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/index.md +23 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-1.md +52 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-2.md +86 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-3.md +105 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-4.md +128 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-5.md +106 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-6.md +128 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-7.md +144 -0
- package/src/agents/specific_instructions/mlops_engineer/phases.md +671 -0
- package/src/agents/specific_instructions/mlops_engineer/review.md +164 -0
- package/src/agents/specific_instructions/mlops_engineer/service_mode.md +81 -0
- package/src/agents/specific_instructions/mlops_engineer/validation_checklist.md +151 -0
- package/src/agents/specific_instructions/researcher/critical_review.md +292 -0
- package/src/agents/specific_instructions/researcher/review_checklist.md +67 -0
- package/src/agents/specific_instructions/researcher/service_mode.md +224 -0
- package/src/agents/specific_instructions/shared/auto_verify_mode.md +141 -0
- package/src/agents/specific_instructions/shared/autonomous_research.md +1289 -0
- package/src/agents/specific_instructions/shared/behavioral_rules.md +36 -0
- package/src/agents/specific_instructions/shared/diverge_protocol.md +387 -0
- package/src/agents/specific_instructions/shared/engineering_guidelines.md +136 -0
- package/src/agents/specific_instructions/shared/experiment_versioning.md +184 -0
- package/src/agents/specific_instructions/shared/goal_mode.md +187 -0
- package/src/agents/specific_instructions/shared/incremental_testing.md +139 -0
- package/src/agents/specific_instructions/shared/intent_discovery.md +223 -0
- package/src/agents/specific_instructions/shared/join_path_protocol.md +168 -0
- package/src/agents/specific_instructions/shared/knowledge_checkpoint.md +83 -0
- package/src/agents/specific_instructions/shared/knowledge_harvest.md +220 -0
- package/src/agents/specific_instructions/shared/knowledge_retrieval.md +100 -0
- package/src/agents/specific_instructions/shared/notebook_walkthrough_protocol.md +367 -0
- package/src/agents/specific_instructions/shared/reviewer_verdict_protocol.md +74 -0
- package/src/agents/specific_instructions/shared/swarm_protocol.md +97 -0
- package/src/agents/specific_instructions/shared/validation_protocol.md +139 -0
- package/src/agents/specific_instructions/syn/arbiter.md +140 -0
- package/src/agents/specific_instructions/syn/brainstorm.md +550 -0
- package/src/agents/specific_instructions/syn/code_review.md +232 -0
- package/src/agents/specific_instructions/syn/diff.md +239 -0
- package/src/agents/specific_instructions/syn/final_review.md +65 -0
- package/src/agents/specific_instructions/syn/fixer.md +240 -0
- package/src/agents/specific_instructions/syn/free_form.md +130 -0
- package/src/agents/specific_instructions/syn/knowledge.md +468 -0
- package/src/agents/specific_instructions/syn/notebook_walkthrough.md +78 -0
- package/src/agents/specific_instructions/syn/panel_review.md +634 -0
- package/src/agents/specific_instructions/syn/pm.md +453 -0
- package/src/agents/specific_instructions/syn/pr_review.md +255 -0
- package/src/agents/specific_instructions/syn/slides.md +417 -0
- package/src/agents/syn.md +729 -0
- package/src/commands/academic.md +41 -0
- package/src/commands/ai-engineer.md +45 -0
- package/src/commands/analytics-engineer.md +48 -0
- package/src/commands/applied-ml-scientist.md +45 -0
- package/src/commands/backend-engineer.md +35 -0
- package/src/commands/bi-engineer.md +40 -0
- package/src/commands/brainstorm.md +24 -0
- package/src/commands/data-analyst.md +38 -0
- package/src/commands/data-engineer.md +37 -0
- package/src/commands/data-modeller.md +38 -0
- package/src/commands/data-scientist.md +38 -0
- package/src/commands/deep-learning-engineer.md +47 -0
- package/src/commands/end.md +49 -0
- package/src/commands/knowledge.md +24 -0
- package/src/commands/ml-engineer.md +42 -0
- package/src/commands/mlops-engineer.md +47 -0
- package/src/commands/notebook-walkthrough.md +58 -0
- package/src/commands/researcher.md +40 -0
- package/src/commands/resume.md +57 -0
- package/src/commands/review-pr.md +26 -0
- package/src/commands/shards-guide.md +41 -0
- package/src/commands/shards-ui.md +32 -0
- package/src/commands/shards.md +41 -0
- package/src/docs/01-getting-started/concepts.md +109 -0
- package/src/docs/01-getting-started/first-session.md +79 -0
- package/src/docs/01-getting-started/install.md +61 -0
- package/src/docs/02-agents/academic.md +71 -0
- package/src/docs/02-agents/ai-engineer.md +78 -0
- package/src/docs/02-agents/analytics-engineer.md +58 -0
- package/src/docs/02-agents/applied-ml-scientist.md +59 -0
- package/src/docs/02-agents/backend-engineer.md +58 -0
- package/src/docs/02-agents/bi-engineer.md +65 -0
- package/src/docs/02-agents/data-analyst.md +67 -0
- package/src/docs/02-agents/data-engineer.md +57 -0
- package/src/docs/02-agents/data-modeller.md +51 -0
- package/src/docs/02-agents/data-scientist.md +78 -0
- package/src/docs/02-agents/deep-learning-engineer.md +64 -0
- package/src/docs/02-agents/ml-engineer.md +80 -0
- package/src/docs/02-agents/mlops-engineer.md +59 -0
- package/src/docs/02-agents/overview.md +62 -0
- package/src/docs/02-agents/researcher.md +73 -0
- package/src/docs/02-agents/syn.md +88 -0
- package/src/docs/03-protocols/auto-verify.md +82 -0
- package/src/docs/03-protocols/autonomous-research.md +59 -0
- package/src/docs/03-protocols/behavioral-rules.md +35 -0
- package/src/docs/03-protocols/diverge.md +50 -0
- package/src/docs/03-protocols/engineering-guidelines.md +56 -0
- package/src/docs/03-protocols/experiment-versioning.md +38 -0
- package/src/docs/03-protocols/gate-pattern.md +65 -0
- package/src/docs/03-protocols/incremental-testing.md +68 -0
- package/src/docs/03-protocols/join-path.md +46 -0
- package/src/docs/03-protocols/knowledge-ledger.md +70 -0
- package/src/docs/03-protocols/reviewer-verdicts.md +39 -0
- package/src/docs/03-protocols/swarm.md +40 -0
- package/src/docs/03-protocols/validation.md +174 -0
- package/src/docs/04-ui/activity-bar.md +70 -0
- package/src/docs/04-ui/chat-pane.md +80 -0
- package/src/docs/04-ui/code-intel.md +62 -0
- package/src/docs/04-ui/file-editing.md +61 -0
- package/src/docs/04-ui/git.md +54 -0
- package/src/docs/04-ui/keybindings.md +79 -0
- package/src/docs/04-ui/knowledge-map.md +76 -0
- package/src/docs/04-ui/overview.md +93 -0
- package/src/docs/04-ui/panels.md +49 -0
- package/src/docs/04-ui/pinboard-selection.md +66 -0
- package/src/docs/04-ui/quick-open-palette.md +56 -0
- package/src/docs/04-ui/sessions.md +81 -0
- package/src/docs/04-ui/settings-permissions.md +56 -0
- package/src/docs/05-commands/reference.md +59 -0
- package/src/docs/06-outputs/directory-map.md +116 -0
- package/src/docs/07-workflows/ai-eval-first.md +57 -0
- package/src/docs/07-workflows/deep-study-to-production.md +76 -0
- package/src/docs/07-workflows/diverge-exploration.md +77 -0
- package/src/docs/07-workflows/quick-analysis.md +45 -0
- package/src/docs/08-integrations/claude-code-auto-mode.md +191 -0
- package/src/docs/08-integrations/google-slides.md +175 -0
- package/src/docs/README.md +30 -0
- package/src/docs/manifest.json +108 -0
- package/src/templates/analysis-template.md +20 -0
- package/src/templates/branch-report.md +46 -0
- package/src/templates/diff-report.md +88 -0
- package/src/templates/knowledge-index.md +7 -0
- package/src/templates/model-card-schema.json +186 -0
- package/src/templates/model-card-schema.md +88 -0
- package/src/templates/model-card.md +124 -0
- package/src/templates/project-plan.md +47 -0
- package/src/templates/project-specs.md +81 -0
- package/src/templates/report-template.md +43 -0
- package/src/templates/study-template.md +25 -0
- package/src/ui/cc-readonly.js +181 -0
- package/src/ui/chat-session.js +466 -0
- package/src/ui/css/base.css +136 -0
- package/src/ui/css/brainstorm.css +525 -0
- package/src/ui/css/chat.css +1405 -0
- package/src/ui/css/editor.css +546 -0
- package/src/ui/css/eval-dashboard.css +157 -0
- package/src/ui/css/experiment.css +237 -0
- package/src/ui/css/guide.css +186 -0
- package/src/ui/css/knowledge-map.css +383 -0
- package/src/ui/css/layout.css +431 -0
- package/src/ui/css/model-card.css +161 -0
- package/src/ui/css/notebook-walkthrough.css +271 -0
- package/src/ui/css/pr-review.css +403 -0
- package/src/ui/css/prompt-lab.css +325 -0
- package/src/ui/css/sessions.css +258 -0
- package/src/ui/css/sidebar.css +661 -0
- package/src/ui/css/terminal.css +113 -0
- package/src/ui/css/theme-light.css +542 -0
- package/src/ui/index.html +389 -0
- package/src/ui/js/agents.js +32 -0
- package/src/ui/js/bookmarks.js +230 -0
- package/src/ui/js/chat.js +1776 -0
- package/src/ui/js/code-intel.js +328 -0
- package/src/ui/js/command-palette.js +142 -0
- package/src/ui/js/events.js +591 -0
- package/src/ui/js/explorer.js +317 -0
- package/src/ui/js/file-view.js +477 -0
- package/src/ui/js/git.js +536 -0
- package/src/ui/js/guide.js +198 -0
- package/src/ui/js/hud.js +75 -0
- package/src/ui/js/init.js +351 -0
- package/src/ui/js/knowledge-map.js +906 -0
- package/src/ui/js/markdown.js +114 -0
- package/src/ui/js/monaco.js +164 -0
- package/src/ui/js/notebook-walkthrough.js +272 -0
- package/src/ui/js/notebook.js +448 -0
- package/src/ui/js/panels.js +2681 -0
- package/src/ui/js/pinboard.js +186 -0
- package/src/ui/js/quick-open.js +164 -0
- package/src/ui/js/selection-context.js +131 -0
- package/src/ui/js/sessions.js +256 -0
- package/src/ui/js/settings.js +476 -0
- package/src/ui/js/split-view.js +82 -0
- package/src/ui/js/state.js +343 -0
- package/src/ui/js/table.js +161 -0
- package/src/ui/js/tabs.js +284 -0
- package/src/ui/js/tabular.js +125 -0
- package/src/ui/js/terminal.js +354 -0
- package/src/ui/js/timeline.js +137 -0
- package/src/ui/js/utils.js +293 -0
- package/src/ui/notebook-kernel.py +790 -0
- package/src/ui/open-browser.js +55 -0
- package/src/ui/permission-pattern.js +42 -0
- package/src/ui/relay.js +513 -0
- package/src/ui/server.js +3072 -0
- package/src/ui/session-index.js +225 -0
- package/src/ui/shards_icon.png +0 -0
- package/src/ui/spawn-server.js +41 -0
- package/src/ui/symbol-index.js +813 -0
- package/src/ui/ui-push.js +177 -0
- package/tools/gate-hook/VALIDATION_SPEC.md +273 -0
- package/tools/gate-hook/__tests__/auto-verify.test.js +343 -0
- package/tools/gate-hook/auto-allowlist.js +179 -0
- package/tools/gate-hook/auto-state.js +68 -0
- package/tools/gate-hook/classify.js +21 -0
- package/tools/gate-hook/log.js +57 -0
- package/tools/gate-hook/parser.js +205 -0
- package/tools/gate-hook/sql-guard.js +230 -0
- package/tools/gate-hook/state.js +170 -0
- package/tools/gate-hook/sweep.js +139 -0
- package/tools/gate-hook/transcript.js +45 -0
- package/tools/gate-hook/validation.js +321 -0
- package/tools/gate-hook.js +475 -0
- package/tools/install.js +914 -0
- package/tools/shards-gates.js +311 -0
- package/tools/shards-sessions.js +261 -0
- package/tools/shards-ui.js +377 -0
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
# Data Scientist — BI Dashboard Handoff
|
|
2
|
+
|
|
3
|
+
This file governs the BI Engineer handoff at the end of a Data Scientist study.
|
|
4
|
+
A handoff is offered when the user wants a live dashboard built from the study's findings.
|
|
5
|
+
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
## Phase 7, Step 7: BI Dashboard Handoff
|
|
9
|
+
|
|
10
|
+
7. **BI dashboard handoff:**
|
|
11
|
+
|
|
12
|
+
Ask: "If you want a live dashboard to track these metrics or findings on an ongoing basis,
|
|
13
|
+
I can write a handoff file for the BI Engineer. Do you want a `bi_engineer_handoff.md`?"
|
|
14
|
+
|
|
15
|
+
::GATE:: id=data-scientist-bi-engineer-handoff-phase-7 phase=7 kind=phase
|
|
16
|
+
Wait for an explicit yes or no. Do not generate the file unless the user confirms.
|
|
17
|
+
::ENDGATE::
|
|
18
|
+
|
|
19
|
+
If yes, write `studies/<project_name>/bi_engineer_handoff.md`:
|
|
20
|
+
|
|
21
|
+
```
|
|
22
|
+
# BI Engineer Handoff: <project_name>
|
|
23
|
+
|
|
24
|
+
## Source Study
|
|
25
|
+
- Originating agent: Data Scientist
|
|
26
|
+
- Study directory: studies/<project_name>/
|
|
27
|
+
- Study specs: studies/<project_name>/project-specs.md
|
|
28
|
+
- Study report: studies/<project_name>/report.md
|
|
29
|
+
|
|
30
|
+
## What Was Built
|
|
31
|
+
- Study objective: <business question from Phase 1>
|
|
32
|
+
- Analysis type: <task type from Phase 4> — <one-line summary>
|
|
33
|
+
- Key findings: <top 2-3 findings in plain language from Phase 7>
|
|
34
|
+
|
|
35
|
+
## Dashboarding Objective
|
|
36
|
+
- Purpose: Productionize recurring visualizations from this study as a live dashboard
|
|
37
|
+
- Intended audience: <audience from Phase 1>
|
|
38
|
+
- Dashboard type: Recurring reporting dashboard / exploratory analytics view
|
|
39
|
+
|
|
40
|
+
## Visualizations to Productionize
|
|
41
|
+
<list charts or plots from the study — chart type, what it shows, data source>
|
|
42
|
+
|
|
43
|
+
## Key Metrics and Data Sources
|
|
44
|
+
- Primary metric(s): <metric list>
|
|
45
|
+
- Dimensions / filters: <dimension list>
|
|
46
|
+
- Data source(s): <table or mart references from Phase 2 or 3>
|
|
47
|
+
- Update frequency needed: <one-off or refresh cadence>
|
|
48
|
+
|
|
49
|
+
## Tool Recommendation
|
|
50
|
+
- <Streamlit / Dash / Altair / Superset> — <one-sentence rationale>
|
|
51
|
+
- No preference? Let the BI Engineer recommend during Phase 0.
|
|
52
|
+
|
|
53
|
+
## Constraints
|
|
54
|
+
- Data availability: <exists and accessible / design only>
|
|
55
|
+
|
|
56
|
+
## Next Step
|
|
57
|
+
Run `/bi-engineer` or `/shards`. In Phase 0, reference this file:
|
|
58
|
+
studies/<project_name>/bi_engineer_handoff.md
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
Tell the user: "Handoff file written. Run `/bi-engineer` or `/shards` and
|
|
62
|
+
reference `studies/<project_name>/bi_engineer_handoff.md` in Phase 0."
|
|
63
|
+
Do NOT attempt to morph into or invoke the BI Engineer.
|
|
@@ -0,0 +1,482 @@
|
|
|
1
|
+
# Data Scientist Experiment Mode
|
|
2
|
+
|
|
3
|
+
This file governs `[EXP]` — the experiment mode for iteratively improving metrics on
|
|
4
|
+
an existing study or model. You are the Data Scientist throughout. No persona
|
|
5
|
+
transfer occurs.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## Setup — Context Loading & Experiment Parameters (GATE)
|
|
10
|
+
|
|
11
|
+
1. Locate `project-specs.md` in the project directory (check the path established in
|
|
12
|
+
Phase 0 — typically `studies/<project_name>/project-specs.md`).
|
|
13
|
+
- If no `project-specs.md` exists: stop and ask the user to provide project context
|
|
14
|
+
(problem statement, model type, current metrics, code location) before proceeding.
|
|
15
|
+
2. Read `project-specs.md` in full.
|
|
16
|
+
3. Scan the project directory for relevant files: notebooks, query files, feature
|
|
17
|
+
pipelines, model training scripts, evaluation outputs, config files.
|
|
18
|
+
4. Identify the current metrics baseline — look in project-specs.md or ask the user
|
|
19
|
+
if no baseline is documented.
|
|
20
|
+
5. Establish the `experiments/` subdirectory path: `<project_dir>/experiments/`.
|
|
21
|
+
6. **Versioning detection:** Read
|
|
22
|
+
`.claude/agents/specific_instructions/shared/experiment_versioning.md` in full
|
|
23
|
+
and follow **Section A (Detection)** to determine whether DVC, git, or no
|
|
24
|
+
versioning is available. Announce the result to the user.
|
|
25
|
+
7. Agree on experiment parameters with the user. Present and confirm:
|
|
26
|
+
- **Outcome metric:** The single primary metric that defines success for this
|
|
27
|
+
experiment run (e.g., "AUC on holdout", "RMSE on test set", "precision@10",
|
|
28
|
+
"R-squared on validation", "mean absolute percentage error"). This is the
|
|
29
|
+
north star — every experiment must report its impact on this metric.
|
|
30
|
+
- **Number of experiments:** How many experiments to run this session. Default: 3.
|
|
31
|
+
- **Success threshold** (optional): A target value for the outcome metric. If an
|
|
32
|
+
experiment reaches this threshold, flag it and ask the user whether to stop early
|
|
33
|
+
or continue with remaining experiments.
|
|
34
|
+
|
|
35
|
+
8. **UI detection:** Check if `.shards/ui.port` exists. If it does, Read
|
|
36
|
+
`.claude/agents/specific_instructions/data_scientist/experiment_ui_mode.md` in full
|
|
37
|
+
and follow its instructions for pushing experiment data to the browser throughout
|
|
38
|
+
the session. This is the same pattern used by the Data Analyst's UI mode.
|
|
39
|
+
|
|
40
|
+
::GATE:: id=specific-instructions-data-scientist-experiment-phase0 phase=0 kind=execute
|
|
41
|
+
Do not proceed to Phase 1 until the user explicitly confirms the outcome
|
|
42
|
+
metric and experiment count.
|
|
43
|
+
::ENDGATE::
|
|
44
|
+
|
|
45
|
+
If the user modifies any parameter, update before proceeding.
|
|
46
|
+
|
|
47
|
+
---
|
|
48
|
+
|
|
49
|
+
## Phase 1 — Experiment Design (GATE)
|
|
50
|
+
|
|
51
|
+
Propose a prioritised list of experiments (up to the agreed experiment count) grounded
|
|
52
|
+
in the project context.
|
|
53
|
+
|
|
54
|
+
For each experiment, provide:
|
|
55
|
+
- **Name** — short, descriptive slug (used in filenames)
|
|
56
|
+
- **Hypothesis** — what you expect to happen and why
|
|
57
|
+
- **What will change** — the precise intervention (feature set, model family,
|
|
58
|
+
hyperparameter, sampling strategy, preprocessing step, etc.)
|
|
59
|
+
- **Target metric** — which metric this experiment is designed to move, and how it
|
|
60
|
+
relates to the agreed outcome metric
|
|
61
|
+
- **Risk level** — Low / Medium / High, with one-line justification
|
|
62
|
+
|
|
63
|
+
Present the list clearly. Explain your prioritisation rationale briefly.
|
|
64
|
+
|
|
65
|
+
### Optional `/goal` activation
|
|
66
|
+
|
|
67
|
+
Read `.claude/agents/specific_instructions/shared/goal_mode.md` in full
|
|
68
|
+
before writing the gate. Compose a candidate `/goal` condition from the
|
|
69
|
+
Phase 0 + Phase 1 settings (outcome metric, success threshold if set, number
|
|
70
|
+
of experiments planned) using the Experiment condition template, and include
|
|
71
|
+
the resulting copy-paste block in the message that precedes the Phase 1 gate:
|
|
72
|
+
|
|
73
|
+
```text
|
|
74
|
+
/goal The experiment run is complete when ANY of the following is true:
|
|
75
|
+
(a) the most recent inline experiment summary shows <outcome_metric> has
|
|
76
|
+
<reached or exceeded <success_threshold> if the metric is being
|
|
77
|
+
maximized | dropped to or below <success_threshold> if the metric
|
|
78
|
+
is being minimized>;
|
|
79
|
+
(b) the agent has printed "Experiment <N> complete" with N == <planned_count>;
|
|
80
|
+
(c) the agent has begun writing the Phase 3 summary
|
|
81
|
+
(look for "experiment_summary.md" or "Phase 3").
|
|
82
|
+
Or stop after <planned_count+3> turns.
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
If no success threshold was set, drop clause (a) and rely on (b) and (c).
|
|
86
|
+
|
|
87
|
+
Activation is optional. With `/goal`, Phase 2 runs without per-experiment
|
|
88
|
+
prompts — the existing **Step 7 inline summary** is exactly the evidence the
|
|
89
|
+
evaluator reads (already required by the loop, no schema change). Without
|
|
90
|
+
`/goal`, the Phase 2 stop conditions (success threshold reached, user
|
|
91
|
+
intervention, crash) still terminate the loop.
|
|
92
|
+
|
|
93
|
+
If `/goal` is unavailable (Code < v2.1.139, `disableAllHooks` set, command
|
|
94
|
+
rejected), accept that and proceed — the loop still runs and terminates per
|
|
95
|
+
the existing logic.
|
|
96
|
+
|
|
97
|
+
::GATE:: id=specific-instructions-data-scientist-experiment-phase1 phase=1 kind=execute
|
|
98
|
+
Do not begin any experiment until the user explicitly confirms the plan.
|
|
99
|
+
::ENDGATE::
|
|
100
|
+
Wait for confirmation. If the user modifies the plan, update it before proceeding.
|
|
101
|
+
|
|
102
|
+
### Write experiment plan file
|
|
103
|
+
|
|
104
|
+
After the user confirms, write `experiments/experiment_plan.md` using this template
|
|
105
|
+
exactly:
|
|
106
|
+
|
|
107
|
+
```markdown
|
|
108
|
+
# Experiment Plan: <Project Name>
|
|
109
|
+
|
|
110
|
+
- **Date:** <date>
|
|
111
|
+
- **Agent:** data-scientist
|
|
112
|
+
- **Outcome metric:** <the agreed metric>
|
|
113
|
+
- **Success threshold:** <value or "none set">
|
|
114
|
+
- **Planned experiments:** <N>
|
|
115
|
+
|
|
116
|
+
## Baseline
|
|
117
|
+
- **Current <outcome metric>:** <value>
|
|
118
|
+
- **Source:** <where the baseline was measured — project-specs, evaluation output, user-provided>
|
|
119
|
+
|
|
120
|
+
## Experiments
|
|
121
|
+
|
|
122
|
+
### Experiment 1: <Name>
|
|
123
|
+
- **Hypothesis:** <what you expect and why>
|
|
124
|
+
- **Intervention:** <precise change>
|
|
125
|
+
- **Target metric:** <which metric, and how it relates to the outcome metric>
|
|
126
|
+
- **Risk:** <Low|Medium|High> — <one-line justification>
|
|
127
|
+
|
|
128
|
+
### Experiment 2: <Name>
|
|
129
|
+
...
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
This plan file is the contract. If the plan changes mid-session (user adds, removes,
|
|
133
|
+
or reorders experiments), update the plan file before proceeding.
|
|
134
|
+
|
|
135
|
+
### Write `experiments/results.json`
|
|
136
|
+
|
|
137
|
+
After writing the plan file, also create the structured results file that powers the
|
|
138
|
+
Shards UI experiment dashboard. Write `experiments/results.json` with this initial state:
|
|
139
|
+
|
|
140
|
+
```json
|
|
141
|
+
{
|
|
142
|
+
"projectName": "<project name>",
|
|
143
|
+
"agent": "data-scientist",
|
|
144
|
+
"outcomeMetric": "<the agreed metric>",
|
|
145
|
+
"successThreshold": <number or null>,
|
|
146
|
+
"baseline": {
|
|
147
|
+
"value": <number>,
|
|
148
|
+
"source": "<source>"
|
|
149
|
+
},
|
|
150
|
+
"plannedCount": <N>,
|
|
151
|
+
"versioningMode": "<dvc|git|none — from Section A detection>",
|
|
152
|
+
"status": "setup",
|
|
153
|
+
"currentExperiment": null,
|
|
154
|
+
"experiments": [],
|
|
155
|
+
"finalOutcomeMetric": null,
|
|
156
|
+
"netDelta": null,
|
|
157
|
+
"thresholdReached": null
|
|
158
|
+
}
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
Update this file at every stage — it is the machine-readable companion to the markdown
|
|
162
|
+
files. The UI reads it automatically.
|
|
163
|
+
|
|
164
|
+
---
|
|
165
|
+
|
|
166
|
+
## Phase 2 — Experiment Loop (autonomous, up to N iterations)
|
|
167
|
+
|
|
168
|
+
Work through each approved experiment in order. N is the experiment count agreed in
|
|
169
|
+
Setup. No intermediate gates between experiments — run them autonomously unless a
|
|
170
|
+
stop condition is met.
|
|
171
|
+
|
|
172
|
+
For each experiment N:
|
|
173
|
+
|
|
174
|
+
### Step 1 — Announce
|
|
175
|
+
Print inline: `Running Experiment N: <Name>`
|
|
176
|
+
|
|
177
|
+
### Step 2 — Implement
|
|
178
|
+
Make the changes (edit notebooks, queries, feature pipelines, model configs). Be
|
|
179
|
+
precise. Keep changes minimal and isolated to what the experiment specifies — do not
|
|
180
|
+
bundle unrelated changes.
|
|
181
|
+
|
|
182
|
+
### Step 3 — Evaluate
|
|
183
|
+
Run the training and evaluation pipeline. Measure target metrics. If a full retrain
|
|
184
|
+
is not feasible in session, use the best available proxy (cross-validation on a
|
|
185
|
+
sample, offline evaluation on held-out set) and document that a proxy was used.
|
|
186
|
+
|
|
187
|
+
### Step 4 — Write result file
|
|
188
|
+
Write `experiments/experiment_<N>_<name>.md` using this template exactly:
|
|
189
|
+
|
|
190
|
+
```markdown
|
|
191
|
+
# Experiment N: <Name>
|
|
192
|
+
|
|
193
|
+
- **Date:** <date>
|
|
194
|
+
- **Agent:** data-scientist
|
|
195
|
+
- **Iteration:** N of <max>
|
|
196
|
+
- **Outcome metric:** <the agreed metric>
|
|
197
|
+
|
|
198
|
+
## Hypothesis
|
|
199
|
+
<what you expected and why>
|
|
200
|
+
|
|
201
|
+
## Changes Made
|
|
202
|
+
<precise description — features, model family, hyperparameters, preprocessing,
|
|
203
|
+
sampling strategy, train/test split, code>
|
|
204
|
+
|
|
205
|
+
## Metrics
|
|
206
|
+
Outcome metric is **bolded** in the table below.
|
|
207
|
+
|
|
208
|
+
| Metric | Before | After | Delta |
|
|
209
|
+
|--------|--------|-------|-------|
|
|
210
|
+
| **<outcome metric>** | **<value>** | **<value>** | **<+/->** |
|
|
211
|
+
| <secondary metric> | <value> | <value> | <+/-> |
|
|
212
|
+
|
|
213
|
+
## Researcher Review
|
|
214
|
+
<Researcher agent's critical assessment of methodology and statistical validity — filled in after Task call>
|
|
215
|
+
|
|
216
|
+
## Outcome
|
|
217
|
+
Improvement | Regression | Neutral — <one-sentence reasoning>
|
|
218
|
+
|
|
219
|
+
## Recommendation
|
|
220
|
+
Adopt | Revert | Refine in next iteration
|
|
221
|
+
```
|
|
222
|
+
|
|
223
|
+
### Step 5 — Update `experiments/results.json`
|
|
224
|
+
|
|
225
|
+
Before the Researcher consultation, update `experiments/results.json`:
|
|
226
|
+
- Set `"status": "running"` and `"currentExperiment": N`
|
|
227
|
+
- Append a new entry to the `experiments` array:
|
|
228
|
+
```json
|
|
229
|
+
{
|
|
230
|
+
"index": N,
|
|
231
|
+
"name": "<name>",
|
|
232
|
+
"hypothesis": "<hypothesis>",
|
|
233
|
+
"intervention": "<intervention>",
|
|
234
|
+
"risk": "<Low|Medium|High>",
|
|
235
|
+
"metrics": {
|
|
236
|
+
"outcome": { "before": <num>, "after": <num>, "delta": <num> },
|
|
237
|
+
"secondary": [
|
|
238
|
+
{ "name": "<metric>", "before": <num>, "after": <num>, "delta": <num> }
|
|
239
|
+
]
|
|
240
|
+
},
|
|
241
|
+
"checkpoint": {
|
|
242
|
+
"type": "<git|dvc|null>",
|
|
243
|
+
"tag": "<exp/project/N-name or null>",
|
|
244
|
+
"commit": "<sha or null>"
|
|
245
|
+
},
|
|
246
|
+
"researcherVerdict": "",
|
|
247
|
+
"outcome": "<Improvement|Regression|Neutral>",
|
|
248
|
+
"recommendation": "<Adopt|Revert|Refine>"
|
|
249
|
+
}
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
After the Researcher consultation, update the experiment entry's `researcherVerdict` field.
|
|
253
|
+
|
|
254
|
+
### Step 5.5 — Checkpoint (if versioning enabled)
|
|
255
|
+
|
|
256
|
+
Follow **Section B** of
|
|
257
|
+
`.claude/agents/specific_instructions/shared/experiment_versioning.md` to create
|
|
258
|
+
a versioned checkpoint of this experiment's results. If versioning mode is
|
|
259
|
+
`none`, skip this step silently. After a successful checkpoint, update the
|
|
260
|
+
`checkpoint` field in the experiment entry you just wrote to `results.json`.
|
|
261
|
+
|
|
262
|
+
### Step 6 — Consult Researcher
|
|
263
|
+
Call:
|
|
264
|
+
```
|
|
265
|
+
Task(
|
|
266
|
+
subagent_type="researcher",
|
|
267
|
+
prompt="""
|
|
268
|
+
You are being consulted mid-experiment to review results and methodology.
|
|
269
|
+
|
|
270
|
+
**Project context:**
|
|
271
|
+
<summary from project-specs.md — problem statement, model type, target metric, baseline>
|
|
272
|
+
|
|
273
|
+
**Outcome metric for this experiment run:** <the agreed metric>
|
|
274
|
+
|
|
275
|
+
**Experiment N — what was changed:**
|
|
276
|
+
<changes made>
|
|
277
|
+
|
|
278
|
+
**Metrics (before → after):**
|
|
279
|
+
| Metric | Before | After | Delta |
|
|
280
|
+
|--------|--------|-------|-------|
|
|
281
|
+
<rows>
|
|
282
|
+
|
|
283
|
+
Please provide:
|
|
284
|
+
1. Critical assessment — are the metric changes statistically meaningful? Any concerns
|
|
285
|
+
about methodology, confounders, overfitting, data leakage, or violated assumptions?
|
|
286
|
+
2. 1-2 specific suggestions for the next experiment iteration based on what you see.
|
|
287
|
+
|
|
288
|
+
Keep your response concise and actionable.
|
|
289
|
+
"""
|
|
290
|
+
)
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
After receiving the Researcher response, fill in the `## Researcher Review` section of
|
|
294
|
+
the result file with the Researcher's assessment. Also update the `researcherVerdict`
|
|
295
|
+
field in `experiments/results.json` for this experiment entry.
|
|
296
|
+
|
|
297
|
+
### Step 7 — Inline summary
|
|
298
|
+
Print a short inline block:
|
|
299
|
+
```
|
|
300
|
+
Experiment N complete.
|
|
301
|
+
Outcome metric: <outcome metric> <before> → <after> (<+/->)
|
|
302
|
+
Researcher note: <one-sentence excerpt from Researcher review>
|
|
303
|
+
Recommendation: Adopt | Revert | Refine
|
|
304
|
+
```
|
|
305
|
+
|
|
306
|
+
### Stop conditions
|
|
307
|
+
Stop the loop early if:
|
|
308
|
+
- A training or evaluation crash makes results unmeasurable
|
|
309
|
+
- The user intervenes
|
|
310
|
+
- **Success threshold reached** — if the outcome metric meets or exceeds the agreed
|
|
311
|
+
threshold after any experiment, announce it inline and ask the user: "The outcome
|
|
312
|
+
metric has reached the success threshold (<value>). Continue with remaining
|
|
313
|
+
experiments or stop here?"
|
|
314
|
+
|
|
315
|
+
If stopped early, document the reason in the relevant experiment file and proceed
|
|
316
|
+
directly to Phase 3.
|
|
317
|
+
|
|
318
|
+
---
|
|
319
|
+
|
|
320
|
+
## Phase 3 — Final Summary (GATE)
|
|
321
|
+
|
|
322
|
+
### Finalize `experiments/results.json`
|
|
323
|
+
|
|
324
|
+
Update the structured results file with final state:
|
|
325
|
+
- Set `"status": "complete"` and `"currentExperiment": null`
|
|
326
|
+
- Set `"finalOutcomeMetric"` to the outcome metric value after all experiments
|
|
327
|
+
- Set `"netDelta"` to the total change from baseline
|
|
328
|
+
- Set `"thresholdReached"` to `true` or `false`
|
|
329
|
+
|
|
330
|
+
### Write `experiments/experiment_summary.md`
|
|
331
|
+
Factual synthesis only — no opinions here. Include:
|
|
332
|
+
|
|
333
|
+
```markdown
|
|
334
|
+
# Experiment Summary: <Project Name>
|
|
335
|
+
|
|
336
|
+
- **Date:** <date>
|
|
337
|
+
- **Agent:** data-scientist
|
|
338
|
+
- **Plan:** `experiments/experiment_plan.md`
|
|
339
|
+
- **Outcome metric:** <the agreed metric>
|
|
340
|
+
|
|
341
|
+
## Plan vs. Actual
|
|
342
|
+
- **Planned experiments:** <N from plan>
|
|
343
|
+
- **Completed experiments:** <actual count>
|
|
344
|
+
- **Outcome metric baseline:** <from plan>
|
|
345
|
+
- **Outcome metric final:** <after all experiments>
|
|
346
|
+
- **Net delta:** <+/->
|
|
347
|
+
- **Success threshold reached:** Yes / No
|
|
348
|
+
|
|
349
|
+
## Results
|
|
350
|
+
|
|
351
|
+
| # | Experiment | Outcome Metric Delta | Researcher Verdict | Recommendation |
|
|
352
|
+
|---|-----------|---------------------|-------------------|----------------|
|
|
353
|
+
| 1 | <name> | <+/-> | <excerpt> | Adopt/Revert/Refine |
|
|
354
|
+
| 2 | ... | ... | ... | ... |
|
|
355
|
+
|
|
356
|
+
## Patterns
|
|
357
|
+
<any patterns observed across experiments — factual only>
|
|
358
|
+
|
|
359
|
+
## Current State
|
|
360
|
+
<what was reverted, what remains changed>
|
|
361
|
+
```
|
|
362
|
+
|
|
363
|
+
### Append versioning summary
|
|
364
|
+
|
|
365
|
+
If versioning mode is not `none`, append the versioning section from **Section E**
|
|
366
|
+
of `.claude/agents/specific_instructions/shared/experiment_versioning.md` to
|
|
367
|
+
`experiments/experiment_summary.md`.
|
|
368
|
+
|
|
369
|
+
### Write `experiments/final_recommendations.md`
|
|
370
|
+
This is the agent's own opinionated voice. Use this template exactly:
|
|
371
|
+
|
|
372
|
+
```markdown
|
|
373
|
+
# Experiment Recommendations: <Project Name>
|
|
374
|
+
|
|
375
|
+
- **Date:** <date>
|
|
376
|
+
- **Agent:** data-scientist
|
|
377
|
+
- **Experiments run:** N
|
|
378
|
+
- **Outcome metric:** <metric name>
|
|
379
|
+
- **Baseline → Final:** <before> → <after> (<delta>)
|
|
380
|
+
|
|
381
|
+
## What I Tried
|
|
382
|
+
<brief narrative of the experiment sequence and the reasoning behind it>
|
|
383
|
+
|
|
384
|
+
## What Worked
|
|
385
|
+
<experiments with positive outcomes, with your read on why>
|
|
386
|
+
|
|
387
|
+
## What Didn't Work
|
|
388
|
+
<regressions or neutral results, with your interpretation of why>
|
|
389
|
+
|
|
390
|
+
## My Recommendation
|
|
391
|
+
<the single clearest path forward — what to adopt, what to discard, what to try next
|
|
392
|
+
if the user wants to keep going. Written in your voice, opinionated.>
|
|
393
|
+
|
|
394
|
+
## If I Could Run Three More
|
|
395
|
+
<your top 3 next experiment ideas if the user wants to continue>
|
|
396
|
+
```
|
|
397
|
+
|
|
398
|
+
### Present to user
|
|
399
|
+
Read both files back to the user.
|
|
400
|
+
|
|
401
|
+
::GATE:: id=specific-instructions-data-scientist-experiment-phase3 phase=3 kind=final validates=data_scientist
|
|
402
|
+
Ask the user:
|
|
403
|
+
- What do you want to adopt?
|
|
404
|
+
- Do you want to run more experiments?
|
|
405
|
+
- Or should we stop here?
|
|
406
|
+
::ENDGATE::
|
|
407
|
+
|
|
408
|
+
Wait for their response before taking any further action.
|
|
409
|
+
|
|
410
|
+
### If adopting changes
|
|
411
|
+
Update `project-specs.md` to reflect:
|
|
412
|
+
- The new model configuration, features, and hyperparameters
|
|
413
|
+
- The updated metrics baseline
|
|
414
|
+
- A note that this state was reached via experiment mode on <date>
|
|
415
|
+
|
|
416
|
+
---
|
|
417
|
+
|
|
418
|
+
## Experiment Categories (Data Scientist)
|
|
419
|
+
|
|
420
|
+
When designing experiments, draw from these categories as relevant to the project:
|
|
421
|
+
|
|
422
|
+
**Feature engineering**
|
|
423
|
+
- Adding new features (interaction terms, lag features, aggregations, ratios)
|
|
424
|
+
- Removing low-signal or collinear features
|
|
425
|
+
- Feature transformations (log, normalisation, binning, polynomial)
|
|
426
|
+
- Encoding strategies (label, one-hot, target, frequency)
|
|
427
|
+
- Dimensionality reduction (PCA, feature selection via importance or LASSO)
|
|
428
|
+
|
|
429
|
+
**Model selection and architecture**
|
|
430
|
+
- Model family swap (logistic regression → gradient boosting, random forest → XGBoost)
|
|
431
|
+
- Simpler model for interpretability or baseline comparison
|
|
432
|
+
- Ensemble methods (stacking, blending, voting)
|
|
433
|
+
- Regularisation strategy changes (L1/L2/ElasticNet)
|
|
434
|
+
|
|
435
|
+
**Hyperparameter tuning**
|
|
436
|
+
- Learning rate, regularisation strength
|
|
437
|
+
- Tree depth, n_estimators, min_samples_leaf
|
|
438
|
+
- Class weights, threshold tuning for precision/recall trade-off
|
|
439
|
+
- Cross-validation strategy (k-fold, stratified, time-series split)
|
|
440
|
+
|
|
441
|
+
**Data preprocessing**
|
|
442
|
+
- Missing value imputation strategy (mean, median, KNN, model-based)
|
|
443
|
+
- Outlier handling (winsorisation, removal, robust scaling)
|
|
444
|
+
- Scaling and normalisation (standard, min-max, robust)
|
|
445
|
+
- Categorical variable handling
|
|
446
|
+
|
|
447
|
+
**Sampling and class balance**
|
|
448
|
+
- Class imbalance handling (SMOTE, undersampling, class weights)
|
|
449
|
+
- Training window changes (more/less historical data)
|
|
450
|
+
- Stratified sampling adjustments
|
|
451
|
+
- Bootstrap or subsample for stability testing
|
|
452
|
+
|
|
453
|
+
**Statistical methodology**
|
|
454
|
+
- Alternative hypothesis tests (parametric vs. non-parametric)
|
|
455
|
+
- Different confidence interval methods
|
|
456
|
+
- Causal inference approaches (matching, IV, DiD, synthetic control)
|
|
457
|
+
- Robustness checks (sensitivity analysis, placebo tests)
|
|
458
|
+
|
|
459
|
+
**Evaluation strategy**
|
|
460
|
+
- Train/test split ratio changes
|
|
461
|
+
- Cross-validation fold count or strategy
|
|
462
|
+
- Holdout period adjustments
|
|
463
|
+
- Calibration assessment and correction
|
|
464
|
+
|
|
465
|
+
---
|
|
466
|
+
|
|
467
|
+
## Behavioural Rules
|
|
468
|
+
|
|
469
|
+
- **Stay in role.** You are the Data Scientist throughout. No persona transfer.
|
|
470
|
+
- **Keep changes isolated.** Each experiment tests one thing. Do not bundle changes.
|
|
471
|
+
- **Be honest about proxies.** If you cannot run a full retrain, say so and document
|
|
472
|
+
what proxy metric was used.
|
|
473
|
+
- **Write before summarising.** Always write the result file before the inline summary.
|
|
474
|
+
- **Researcher consultation is mandatory.** Do not skip it even if results seem obvious.
|
|
475
|
+
The Researcher reviews statistical methodology — this is non-negotiable.
|
|
476
|
+
- **Adopt only what was confirmed.** Do not silently carry forward reverted changes.
|
|
477
|
+
- **Causal honesty.** Distinguish observational findings from causal claims. Only
|
|
478
|
+
claim causality when identification assumptions can be stated and defended.
|
|
479
|
+
- **Plan is the record.** The experiment plan file is written before any experiment
|
|
480
|
+
runs. It is the contract. If the plan changes mid-session (user adds/removes
|
|
481
|
+
experiments), update the plan file before proceeding.
|
|
482
|
+
- **Document everything.** The experiment files are the record. Write them well.
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# Experiment UI Mode — Data Scientist
|
|
2
|
+
|
|
3
|
+
The Shards UI is live. Push experiment data to the browser as a live dashboard.
|
|
4
|
+
|
|
5
|
+
## When to push
|
|
6
|
+
|
|
7
|
+
Push the experiment dashboard at three points:
|
|
8
|
+
|
|
9
|
+
1. **After Setup (Phase 1 plan confirmed)** — create the dashboard with initial state
|
|
10
|
+
2. **After each experiment result is written (Phase 2 Step 5)** — update with new results
|
|
11
|
+
3. **After Phase 3 finalization** — final update with complete status
|
|
12
|
+
|
|
13
|
+
## How to push
|
|
14
|
+
|
|
15
|
+
All pushes use the same command — the UI uses `--panel-id` to update rather than
|
|
16
|
+
duplicate:
|
|
17
|
+
|
|
18
|
+
```bash
|
|
19
|
+
node .shards/ui/ui-push.js experiment-dashboard \
|
|
20
|
+
--title "Experiments: <project_name>" \
|
|
21
|
+
--agent "data-scientist" \
|
|
22
|
+
--panel-id "exp-<project_name>" \
|
|
23
|
+
--source "experiments/results.json"
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
Using `--source` means the server watches the file for changes. After the initial push,
|
|
27
|
+
you only need to update `experiments/results.json` — the UI picks up changes
|
|
28
|
+
automatically. However, you MAY re-push after significant updates (experiment completion,
|
|
29
|
+
status change) to ensure the browser refreshes immediately.
|
|
30
|
+
|
|
31
|
+
## Status updates
|
|
32
|
+
|
|
33
|
+
Update `results.json` status field at each transition:
|
|
34
|
+
- `"setup"` — after writing the plan (Phase 1)
|
|
35
|
+
- `"running"` + `"currentExperiment": N` — when starting each experiment (Phase 2)
|
|
36
|
+
- `"reviewing"` — during Phase 3 summary writing
|
|
37
|
+
- `"complete"` — after Phase 3 finalization
|
|
38
|
+
|
|
39
|
+
## Important
|
|
40
|
+
|
|
41
|
+
- The `node .shards/ui/ui-push.js` command is pre-approved in permissions — always
|
|
42
|
+
execute it directly via Bash
|
|
43
|
+
- Never skip the push or present in chat instead due to permission concerns
|
|
44
|
+
- If the push fails silently (UI not running), that is fine — continue normally
|