@proflandrigan/shards 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +475 -0
- package/package.json +37 -0
- package/src/agents/academic.md +276 -0
- package/src/agents/ai-engineer.md +377 -0
- package/src/agents/analytics-engineer.md +364 -0
- package/src/agents/applied-ml-scientist.md +410 -0
- package/src/agents/backend-engineer.md +255 -0
- package/src/agents/bi-engineer.md +333 -0
- package/src/agents/data-analyst.md +343 -0
- package/src/agents/data-engineer.md +260 -0
- package/src/agents/data-modeller.md +386 -0
- package/src/agents/data-scientist.md +366 -0
- package/src/agents/deep-learning-engineer.md +389 -0
- package/src/agents/ml-engineer.md +424 -0
- package/src/agents/mlops-engineer.md +339 -0
- package/src/agents/researcher.md +187 -0
- package/src/agents/specific_instructions/academic/critical_review.md +263 -0
- package/src/agents/specific_instructions/academic/report.md +113 -0
- package/src/agents/specific_instructions/ai_engineer/advise.md +162 -0
- package/src/agents/specific_instructions/ai_engineer/bi_engineer_handoff.md +86 -0
- package/src/agents/specific_instructions/ai_engineer/experiment.md +471 -0
- package/src/agents/specific_instructions/ai_engineer/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/ai_engineer/phases/index.md +45 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-1.md +55 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-2.md +86 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-3.md +96 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-4.md +138 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-5.md +157 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-6.md +196 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-7.md +313 -0
- package/src/agents/specific_instructions/ai_engineer/phases.md +1011 -0
- package/src/agents/specific_instructions/ai_engineer/prompt_lab.md +161 -0
- package/src/agents/specific_instructions/ai_engineer/prompt_lab_ui_mode.md +28 -0
- package/src/agents/specific_instructions/ai_engineer/research.md +393 -0
- package/src/agents/specific_instructions/ai_engineer/research_ui_mode.md +66 -0
- package/src/agents/specific_instructions/ai_engineer/review.md +159 -0
- package/src/agents/specific_instructions/ai_engineer/validation_checklist.md +182 -0
- package/src/agents/specific_instructions/analytics_engineer/advise.md +155 -0
- package/src/agents/specific_instructions/analytics_engineer/bi_engineer_handoff.md +91 -0
- package/src/agents/specific_instructions/analytics_engineer/data_analyst_handoff.md +84 -0
- package/src/agents/specific_instructions/analytics_engineer/deep_phases.md +818 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/index.md +24 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-1.md +77 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-2.md +106 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-3.md +93 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-4.md +79 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-5.md +61 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-6.md +45 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-7.md +235 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-8.md +221 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-1.md +47 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-2.md +78 -0
- package/src/agents/specific_instructions/analytics_engineer/quick_phases.md +112 -0
- package/src/agents/specific_instructions/analytics_engineer/review.md +167 -0
- package/src/agents/specific_instructions/analytics_engineer/service_mode.md +369 -0
- package/src/agents/specific_instructions/analytics_engineer/ui_mode.md +45 -0
- package/src/agents/specific_instructions/analytics_engineer/update.md +162 -0
- package/src/agents/specific_instructions/analytics_engineer/validation_checklist.md +121 -0
- package/src/agents/specific_instructions/applied_ml_scientist/advise.md +143 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/index.md +21 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-1.md +51 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-2.md +66 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-3.md +113 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-4.md +104 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-5.md +156 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases.md +428 -0
- package/src/agents/specific_instructions/applied_ml_scientist/research.md +379 -0
- package/src/agents/specific_instructions/applied_ml_scientist/review.md +142 -0
- package/src/agents/specific_instructions/applied_ml_scientist/validation_checklist.md +136 -0
- package/src/agents/specific_instructions/backend_engineer/clean.md +149 -0
- package/src/agents/specific_instructions/backend_engineer/review.md +91 -0
- package/src/agents/specific_instructions/backend_engineer/review_checklist.md +54 -0
- package/src/agents/specific_instructions/backend_engineer/service_mode.md +67 -0
- package/src/agents/specific_instructions/bi_engineer/advise.md +137 -0
- package/src/agents/specific_instructions/bi_engineer/data_analyst_handoff.md +77 -0
- package/src/agents/specific_instructions/bi_engineer/incoming_handoff.md +45 -0
- package/src/agents/specific_instructions/bi_engineer/phases/index.md +20 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-1.md +164 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-2.md +92 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-3.md +121 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-4.md +106 -0
- package/src/agents/specific_instructions/bi_engineer/phases.md +451 -0
- package/src/agents/specific_instructions/bi_engineer/review.md +166 -0
- package/src/agents/specific_instructions/bi_engineer/update.md +147 -0
- package/src/agents/specific_instructions/bi_engineer/validation_checklist.md +124 -0
- package/src/agents/specific_instructions/data_analyst/advise.md +138 -0
- package/src/agents/specific_instructions/data_analyst/explain.md +221 -0
- package/src/agents/specific_instructions/data_analyst/incoming_handoff.md +40 -0
- package/src/agents/specific_instructions/data_analyst/phases/index.md +20 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-1.md +159 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-2.md +112 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-3.md +265 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-4.md +100 -0
- package/src/agents/specific_instructions/data_analyst/phases.md +501 -0
- package/src/agents/specific_instructions/data_analyst/review.md +138 -0
- package/src/agents/specific_instructions/data_analyst/ui_mode.md +26 -0
- package/src/agents/specific_instructions/data_analyst/update.md +144 -0
- package/src/agents/specific_instructions/data_analyst/validation_checklist.md +95 -0
- package/src/agents/specific_instructions/data_engineer/advise.md +137 -0
- package/src/agents/specific_instructions/data_engineer/phases.md +466 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/index.md +23 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-1.md +49 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-2.md +93 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-3.md +55 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-4.md +48 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-5.md +40 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-6.md +102 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-7.md +87 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/phase-1.md +45 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/phase-2.md +54 -0
- package/src/agents/specific_instructions/data_engineer/review.md +135 -0
- package/src/agents/specific_instructions/data_engineer/validation_checklist.md +136 -0
- package/src/agents/specific_instructions/data_modeller/advise.md +137 -0
- package/src/agents/specific_instructions/data_modeller/phases.md +581 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/index.md +23 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-1.md +52 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-2.md +113 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-3.md +47 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-4.md +51 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-5.md +45 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-6.md +105 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-7.md +136 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/phase-1.md +47 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/phase-2.md +65 -0
- package/src/agents/specific_instructions/data_modeller/review.md +141 -0
- package/src/agents/specific_instructions/data_modeller/service_mode.md +218 -0
- package/src/agents/specific_instructions/data_modeller/validation_checklist.md +125 -0
- package/src/agents/specific_instructions/data_scientist/advise.md +158 -0
- package/src/agents/specific_instructions/data_scientist/bi_engineer_handoff.md +63 -0
- package/src/agents/specific_instructions/data_scientist/experiment.md +482 -0
- package/src/agents/specific_instructions/data_scientist/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/data_scientist/explain.md +247 -0
- package/src/agents/specific_instructions/data_scientist/greenfield_data.md +35 -0
- package/src/agents/specific_instructions/data_scientist/ml_engineer_handoff.md +52 -0
- package/src/agents/specific_instructions/data_scientist/notebook_walkthrough.md +76 -0
- package/src/agents/specific_instructions/data_scientist/phases/index.md +24 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-1.md +45 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-2.md +67 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-3.md +89 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-4.md +143 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-5.md +71 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-6.md +239 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-7.md +207 -0
- package/src/agents/specific_instructions/data_scientist/phases.md +651 -0
- package/src/agents/specific_instructions/data_scientist/research.md +345 -0
- package/src/agents/specific_instructions/data_scientist/research_ui_mode.md +52 -0
- package/src/agents/specific_instructions/data_scientist/review.md +136 -0
- package/src/agents/specific_instructions/data_scientist/service_mode.md +247 -0
- package/src/agents/specific_instructions/data_scientist/validation_checklist.md +183 -0
- package/src/agents/specific_instructions/deep_learning_engineer/advise.md +145 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/index.md +21 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-1.md +74 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-2.md +98 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-3.md +76 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-4.md +128 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-5.md +292 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases.md +567 -0
- package/src/agents/specific_instructions/deep_learning_engineer/research.md +389 -0
- package/src/agents/specific_instructions/deep_learning_engineer/review.md +155 -0
- package/src/agents/specific_instructions/deep_learning_engineer/validation_checklist.md +147 -0
- package/src/agents/specific_instructions/ml_engineer/advise.md +174 -0
- package/src/agents/specific_instructions/ml_engineer/bi_engineer_handoff.md +71 -0
- package/src/agents/specific_instructions/ml_engineer/experiment.md +474 -0
- package/src/agents/specific_instructions/ml_engineer/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/ml_engineer/notebook_walkthrough.md +75 -0
- package/src/agents/specific_instructions/ml_engineer/phases/index.md +25 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-1.md +49 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-2.md +75 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-3.md +124 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-4.md +279 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-5.md +160 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-6-5.md +170 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-6.md +295 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-7.md +337 -0
- package/src/agents/specific_instructions/ml_engineer/phases.md +1068 -0
- package/src/agents/specific_instructions/ml_engineer/research.md +437 -0
- package/src/agents/specific_instructions/ml_engineer/research_ui_mode.md +71 -0
- package/src/agents/specific_instructions/ml_engineer/review.md +187 -0
- package/src/agents/specific_instructions/ml_engineer/service_mode.md +273 -0
- package/src/agents/specific_instructions/ml_engineer/validation_checklist.md +185 -0
- package/src/agents/specific_instructions/mlops_engineer/advise.md +139 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/index.md +23 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-1.md +52 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-2.md +86 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-3.md +105 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-4.md +128 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-5.md +106 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-6.md +128 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-7.md +144 -0
- package/src/agents/specific_instructions/mlops_engineer/phases.md +671 -0
- package/src/agents/specific_instructions/mlops_engineer/review.md +164 -0
- package/src/agents/specific_instructions/mlops_engineer/service_mode.md +81 -0
- package/src/agents/specific_instructions/mlops_engineer/validation_checklist.md +151 -0
- package/src/agents/specific_instructions/researcher/critical_review.md +292 -0
- package/src/agents/specific_instructions/researcher/review_checklist.md +67 -0
- package/src/agents/specific_instructions/researcher/service_mode.md +224 -0
- package/src/agents/specific_instructions/shared/auto_verify_mode.md +141 -0
- package/src/agents/specific_instructions/shared/autonomous_research.md +1289 -0
- package/src/agents/specific_instructions/shared/behavioral_rules.md +36 -0
- package/src/agents/specific_instructions/shared/diverge_protocol.md +387 -0
- package/src/agents/specific_instructions/shared/engineering_guidelines.md +136 -0
- package/src/agents/specific_instructions/shared/experiment_versioning.md +184 -0
- package/src/agents/specific_instructions/shared/goal_mode.md +187 -0
- package/src/agents/specific_instructions/shared/incremental_testing.md +139 -0
- package/src/agents/specific_instructions/shared/intent_discovery.md +223 -0
- package/src/agents/specific_instructions/shared/join_path_protocol.md +168 -0
- package/src/agents/specific_instructions/shared/knowledge_checkpoint.md +83 -0
- package/src/agents/specific_instructions/shared/knowledge_harvest.md +220 -0
- package/src/agents/specific_instructions/shared/knowledge_retrieval.md +100 -0
- package/src/agents/specific_instructions/shared/notebook_walkthrough_protocol.md +367 -0
- package/src/agents/specific_instructions/shared/reviewer_verdict_protocol.md +74 -0
- package/src/agents/specific_instructions/shared/swarm_protocol.md +97 -0
- package/src/agents/specific_instructions/shared/validation_protocol.md +139 -0
- package/src/agents/specific_instructions/syn/arbiter.md +140 -0
- package/src/agents/specific_instructions/syn/brainstorm.md +550 -0
- package/src/agents/specific_instructions/syn/code_review.md +232 -0
- package/src/agents/specific_instructions/syn/diff.md +239 -0
- package/src/agents/specific_instructions/syn/final_review.md +65 -0
- package/src/agents/specific_instructions/syn/fixer.md +240 -0
- package/src/agents/specific_instructions/syn/free_form.md +130 -0
- package/src/agents/specific_instructions/syn/knowledge.md +468 -0
- package/src/agents/specific_instructions/syn/notebook_walkthrough.md +78 -0
- package/src/agents/specific_instructions/syn/panel_review.md +634 -0
- package/src/agents/specific_instructions/syn/pm.md +453 -0
- package/src/agents/specific_instructions/syn/pr_review.md +255 -0
- package/src/agents/specific_instructions/syn/slides.md +417 -0
- package/src/agents/syn.md +729 -0
- package/src/commands/academic.md +41 -0
- package/src/commands/ai-engineer.md +45 -0
- package/src/commands/analytics-engineer.md +48 -0
- package/src/commands/applied-ml-scientist.md +45 -0
- package/src/commands/backend-engineer.md +35 -0
- package/src/commands/bi-engineer.md +40 -0
- package/src/commands/brainstorm.md +24 -0
- package/src/commands/data-analyst.md +38 -0
- package/src/commands/data-engineer.md +37 -0
- package/src/commands/data-modeller.md +38 -0
- package/src/commands/data-scientist.md +38 -0
- package/src/commands/deep-learning-engineer.md +47 -0
- package/src/commands/end.md +49 -0
- package/src/commands/knowledge.md +24 -0
- package/src/commands/ml-engineer.md +42 -0
- package/src/commands/mlops-engineer.md +47 -0
- package/src/commands/notebook-walkthrough.md +58 -0
- package/src/commands/researcher.md +40 -0
- package/src/commands/resume.md +57 -0
- package/src/commands/review-pr.md +26 -0
- package/src/commands/shards-guide.md +41 -0
- package/src/commands/shards-ui.md +32 -0
- package/src/commands/shards.md +41 -0
- package/src/docs/01-getting-started/concepts.md +109 -0
- package/src/docs/01-getting-started/first-session.md +79 -0
- package/src/docs/01-getting-started/install.md +61 -0
- package/src/docs/02-agents/academic.md +71 -0
- package/src/docs/02-agents/ai-engineer.md +78 -0
- package/src/docs/02-agents/analytics-engineer.md +58 -0
- package/src/docs/02-agents/applied-ml-scientist.md +59 -0
- package/src/docs/02-agents/backend-engineer.md +58 -0
- package/src/docs/02-agents/bi-engineer.md +65 -0
- package/src/docs/02-agents/data-analyst.md +67 -0
- package/src/docs/02-agents/data-engineer.md +57 -0
- package/src/docs/02-agents/data-modeller.md +51 -0
- package/src/docs/02-agents/data-scientist.md +78 -0
- package/src/docs/02-agents/deep-learning-engineer.md +64 -0
- package/src/docs/02-agents/ml-engineer.md +80 -0
- package/src/docs/02-agents/mlops-engineer.md +59 -0
- package/src/docs/02-agents/overview.md +62 -0
- package/src/docs/02-agents/researcher.md +73 -0
- package/src/docs/02-agents/syn.md +88 -0
- package/src/docs/03-protocols/auto-verify.md +82 -0
- package/src/docs/03-protocols/autonomous-research.md +59 -0
- package/src/docs/03-protocols/behavioral-rules.md +35 -0
- package/src/docs/03-protocols/diverge.md +50 -0
- package/src/docs/03-protocols/engineering-guidelines.md +56 -0
- package/src/docs/03-protocols/experiment-versioning.md +38 -0
- package/src/docs/03-protocols/gate-pattern.md +65 -0
- package/src/docs/03-protocols/incremental-testing.md +68 -0
- package/src/docs/03-protocols/join-path.md +46 -0
- package/src/docs/03-protocols/knowledge-ledger.md +70 -0
- package/src/docs/03-protocols/reviewer-verdicts.md +39 -0
- package/src/docs/03-protocols/swarm.md +40 -0
- package/src/docs/03-protocols/validation.md +174 -0
- package/src/docs/04-ui/activity-bar.md +70 -0
- package/src/docs/04-ui/chat-pane.md +80 -0
- package/src/docs/04-ui/code-intel.md +62 -0
- package/src/docs/04-ui/file-editing.md +61 -0
- package/src/docs/04-ui/git.md +54 -0
- package/src/docs/04-ui/keybindings.md +79 -0
- package/src/docs/04-ui/knowledge-map.md +76 -0
- package/src/docs/04-ui/overview.md +93 -0
- package/src/docs/04-ui/panels.md +49 -0
- package/src/docs/04-ui/pinboard-selection.md +66 -0
- package/src/docs/04-ui/quick-open-palette.md +56 -0
- package/src/docs/04-ui/sessions.md +81 -0
- package/src/docs/04-ui/settings-permissions.md +56 -0
- package/src/docs/05-commands/reference.md +59 -0
- package/src/docs/06-outputs/directory-map.md +116 -0
- package/src/docs/07-workflows/ai-eval-first.md +57 -0
- package/src/docs/07-workflows/deep-study-to-production.md +76 -0
- package/src/docs/07-workflows/diverge-exploration.md +77 -0
- package/src/docs/07-workflows/quick-analysis.md +45 -0
- package/src/docs/08-integrations/claude-code-auto-mode.md +191 -0
- package/src/docs/08-integrations/google-slides.md +175 -0
- package/src/docs/README.md +30 -0
- package/src/docs/manifest.json +108 -0
- package/src/templates/analysis-template.md +20 -0
- package/src/templates/branch-report.md +46 -0
- package/src/templates/diff-report.md +88 -0
- package/src/templates/knowledge-index.md +7 -0
- package/src/templates/model-card-schema.json +186 -0
- package/src/templates/model-card-schema.md +88 -0
- package/src/templates/model-card.md +124 -0
- package/src/templates/project-plan.md +47 -0
- package/src/templates/project-specs.md +81 -0
- package/src/templates/report-template.md +43 -0
- package/src/templates/study-template.md +25 -0
- package/src/ui/cc-readonly.js +181 -0
- package/src/ui/chat-session.js +466 -0
- package/src/ui/css/base.css +136 -0
- package/src/ui/css/brainstorm.css +525 -0
- package/src/ui/css/chat.css +1405 -0
- package/src/ui/css/editor.css +546 -0
- package/src/ui/css/eval-dashboard.css +157 -0
- package/src/ui/css/experiment.css +237 -0
- package/src/ui/css/guide.css +186 -0
- package/src/ui/css/knowledge-map.css +383 -0
- package/src/ui/css/layout.css +431 -0
- package/src/ui/css/model-card.css +161 -0
- package/src/ui/css/notebook-walkthrough.css +271 -0
- package/src/ui/css/pr-review.css +403 -0
- package/src/ui/css/prompt-lab.css +325 -0
- package/src/ui/css/sessions.css +258 -0
- package/src/ui/css/sidebar.css +661 -0
- package/src/ui/css/terminal.css +113 -0
- package/src/ui/css/theme-light.css +542 -0
- package/src/ui/index.html +389 -0
- package/src/ui/js/agents.js +32 -0
- package/src/ui/js/bookmarks.js +230 -0
- package/src/ui/js/chat.js +1776 -0
- package/src/ui/js/code-intel.js +328 -0
- package/src/ui/js/command-palette.js +142 -0
- package/src/ui/js/events.js +591 -0
- package/src/ui/js/explorer.js +317 -0
- package/src/ui/js/file-view.js +477 -0
- package/src/ui/js/git.js +536 -0
- package/src/ui/js/guide.js +198 -0
- package/src/ui/js/hud.js +75 -0
- package/src/ui/js/init.js +351 -0
- package/src/ui/js/knowledge-map.js +906 -0
- package/src/ui/js/markdown.js +114 -0
- package/src/ui/js/monaco.js +164 -0
- package/src/ui/js/notebook-walkthrough.js +272 -0
- package/src/ui/js/notebook.js +448 -0
- package/src/ui/js/panels.js +2681 -0
- package/src/ui/js/pinboard.js +186 -0
- package/src/ui/js/quick-open.js +164 -0
- package/src/ui/js/selection-context.js +131 -0
- package/src/ui/js/sessions.js +256 -0
- package/src/ui/js/settings.js +476 -0
- package/src/ui/js/split-view.js +82 -0
- package/src/ui/js/state.js +343 -0
- package/src/ui/js/table.js +161 -0
- package/src/ui/js/tabs.js +284 -0
- package/src/ui/js/tabular.js +125 -0
- package/src/ui/js/terminal.js +354 -0
- package/src/ui/js/timeline.js +137 -0
- package/src/ui/js/utils.js +293 -0
- package/src/ui/notebook-kernel.py +790 -0
- package/src/ui/open-browser.js +55 -0
- package/src/ui/permission-pattern.js +42 -0
- package/src/ui/relay.js +513 -0
- package/src/ui/server.js +3072 -0
- package/src/ui/session-index.js +225 -0
- package/src/ui/shards_icon.png +0 -0
- package/src/ui/spawn-server.js +41 -0
- package/src/ui/symbol-index.js +813 -0
- package/src/ui/ui-push.js +177 -0
- package/tools/gate-hook/VALIDATION_SPEC.md +273 -0
- package/tools/gate-hook/__tests__/auto-verify.test.js +343 -0
- package/tools/gate-hook/auto-allowlist.js +179 -0
- package/tools/gate-hook/auto-state.js +68 -0
- package/tools/gate-hook/classify.js +21 -0
- package/tools/gate-hook/log.js +57 -0
- package/tools/gate-hook/parser.js +205 -0
- package/tools/gate-hook/sql-guard.js +230 -0
- package/tools/gate-hook/state.js +170 -0
- package/tools/gate-hook/sweep.js +139 -0
- package/tools/gate-hook/transcript.js +45 -0
- package/tools/gate-hook/validation.js +321 -0
- package/tools/gate-hook.js +475 -0
- package/tools/install.js +914 -0
- package/tools/shards-gates.js +311 -0
- package/tools/shards-sessions.js +261 -0
- package/tools/shards-ui.js +377 -0
|
@@ -0,0 +1,86 @@
|
|
|
1
|
+
> **Previous:** phase-1.md confirmed
|
|
2
|
+
> **Next:** phase-3.md (read only after this phase's gate is confirmed)
|
|
3
|
+
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## Phase 2 — Infrastructure Assessment
|
|
7
|
+
|
|
8
|
+
Goal: Understand existing infrastructure and constraints before designing anything.
|
|
9
|
+
|
|
10
|
+
Ask about:
|
|
11
|
+
- **Existing ML infrastructure:** model registry, feature store, serving layer,
|
|
12
|
+
training orchestrator, experiment tracker — what exists, what doesn't
|
|
13
|
+
- **Cloud services available:** Which managed services can we use? Any
|
|
14
|
+
organizational restrictions?
|
|
15
|
+
- **Compliance / security constraints:** Data residency requirements, VPC
|
|
16
|
+
isolation, IAM constraints, PII handling, audit logging requirements
|
|
17
|
+
- **Team capabilities:** What tooling does the team already know and operate?
|
|
18
|
+
(Important — the right tool for a team that knows Kubernetes is different from
|
|
19
|
+
the right tool for a team that doesn't)
|
|
20
|
+
- **Data pipeline integration:** How do features arrive at training time?
|
|
21
|
+
At serving time? What's the existing data infrastructure?
|
|
22
|
+
- **Existing monitoring:** Any existing observability stack (Prometheus,
|
|
23
|
+
Datadog, CloudWatch, etc.) that ML monitoring should integrate with?
|
|
24
|
+
|
|
25
|
+
**Consult the ML Engineer** for model architecture constraints that affect serving:
|
|
26
|
+
|
|
27
|
+
Tell the user: "Getting the ML Engineer in here — I need to know what the model actually requires before I design serving infrastructure around assumptions."
|
|
28
|
+
|
|
29
|
+
```
|
|
30
|
+
Task(
|
|
31
|
+
subagent_type="ml-engineer",
|
|
32
|
+
description="Review model architecture constraints for MLOps serving design",
|
|
33
|
+
prompt="I am the MLOps Engineer shard scoping an MLOps project for: [project description].
|
|
34
|
+
I need to understand the model architecture constraints that affect my serving and
|
|
35
|
+
infrastructure design. Please tell me:
|
|
36
|
+
1. What is the model framework and format (scikit-learn, XGBoost, PyTorch, TensorFlow, etc.)?
|
|
37
|
+
2. What are the model size and memory requirements (serialized size, memory at inference)?
|
|
38
|
+
3. What are the serving-time feature requirements (features needed at inference, latency sensitivity)?
|
|
39
|
+
4. Does the model support batch inference, online inference, or both?
|
|
40
|
+
5. Are there GPU requirements for inference?
|
|
41
|
+
6. Any known serving constraints or failure modes for this model type?
|
|
42
|
+
7. What's the expected retraining cadence and artifact size?
|
|
43
|
+
Keep the response focused on serving constraints — I'll handle the operational design."
|
|
44
|
+
)
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
### Document Phase 2
|
|
48
|
+
|
|
49
|
+
```markdown
|
|
50
|
+
---
|
|
51
|
+
|
|
52
|
+
## Phase 2: Infrastructure Assessment (MLOps Engineer)
|
|
53
|
+
- **Existing ML infrastructure:**
|
|
54
|
+
- Model registry: <exists — describe | needs setup | N/A>
|
|
55
|
+
- Feature store: <exists — describe | needs setup | N/A>
|
|
56
|
+
- Serving layer: <exists — describe | needs design>
|
|
57
|
+
- Training orchestrator: <Airflow | Kubeflow | SageMaker Pipelines | Vertex AI Pipelines | none>
|
|
58
|
+
- Experiment tracker: <MLflow | W&B | SageMaker Experiments | none>
|
|
59
|
+
- **Cloud services available:** <list of relevant managed services>
|
|
60
|
+
- **Compliance / security:**
|
|
61
|
+
- Data residency: <constraints or "none">
|
|
62
|
+
- VPC / network: <constraints or "none">
|
|
63
|
+
- IAM: <constraints or "none">
|
|
64
|
+
- Audit logging: <required | not required>
|
|
65
|
+
- **Team capabilities:** <what tooling they know and operate>
|
|
66
|
+
- **Data pipeline integration:**
|
|
67
|
+
- Training-time features: <how they arrive>
|
|
68
|
+
- Serving-time features: <how they arrive>
|
|
69
|
+
- Feature freshness: <SLA>
|
|
70
|
+
- **Existing observability stack:** <tools or "none">
|
|
71
|
+
- **ML Engineer consultation:**
|
|
72
|
+
- Model framework: <framework and format>
|
|
73
|
+
- Model size: ~<X>MB serialized, ~<X>MB at inference
|
|
74
|
+
- GPU required for inference: Yes | No
|
|
75
|
+
- Serving constraints: <summary of findings>
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
::GATE:: id=mlops-engineer-phase-2 phase=2 kind=phase
|
|
79
|
+
Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
|
|
80
|
+
::ENDGATE::
|
|
81
|
+
|
|
82
|
+
---
|
|
83
|
+
|
|
84
|
+
## When this gate is confirmed
|
|
85
|
+
|
|
86
|
+
Read `.claude/agents/specific_instructions/mlops_engineer/phases/phase-3.md` in full and follow its instructions starting from Phase 3. Do not pre-read further phase files.
|
|
@@ -0,0 +1,105 @@
|
|
|
1
|
+
> **Previous:** phase-2.md confirmed
|
|
2
|
+
> **Next:** phase-4.md (read only after this phase's gate is confirmed)
|
|
3
|
+
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## Phase 3 — Deployment Design
|
|
7
|
+
|
|
8
|
+
Goal: Design how the model is packaged, served, and versioned.
|
|
9
|
+
|
|
10
|
+
Design decisions to make:
|
|
11
|
+
|
|
12
|
+
**Serving framework selection:**
|
|
13
|
+
Choose based on model type, team capabilities, and cloud:
|
|
14
|
+
- **BentoML** — flexible, framework-agnostic, supports custom pre/post-processing,
|
|
15
|
+
good for teams wanting portability. Operational overhead.
|
|
16
|
+
- **TorchServe** — PyTorch-native, well-integrated with PyTorch ecosystem.
|
|
17
|
+
Less flexible for non-PyTorch models.
|
|
18
|
+
- **NVIDIA Triton Inference Server** — best for GPU inference, multi-model serving,
|
|
19
|
+
high-throughput. Significant operational overhead.
|
|
20
|
+
- **FastAPI + custom** — maximum flexibility, maximum operational overhead.
|
|
21
|
+
Good for simple models, bad for complex serving requirements.
|
|
22
|
+
- **SageMaker Endpoints** — fully managed on AWS, excellent scaling, high cost,
|
|
23
|
+
vendor lock-in. Right choice if team lives in AWS.
|
|
24
|
+
- **Vertex AI Endpoints** — fully managed on GCP, excellent scaling, high cost,
|
|
25
|
+
vendor lock-in. Right choice if team lives in GCP.
|
|
26
|
+
- **Kubernetes + custom** — maximum portability, maximum operational complexity.
|
|
27
|
+
|
|
28
|
+
**Model packaging strategy:**
|
|
29
|
+
- Docker container with model artifacts
|
|
30
|
+
- BentoML Service (`.bento` archive)
|
|
31
|
+
- ONNX export (framework-agnostic, good for latency)
|
|
32
|
+
- TorchScript (PyTorch inference without Python interpreter)
|
|
33
|
+
- MLflow Model (standard format, registry-compatible)
|
|
34
|
+
|
|
35
|
+
**Endpoint design:**
|
|
36
|
+
- REST vs. gRPC (gRPC for high-throughput, latency-sensitive; REST for simplicity)
|
|
37
|
+
- Real-time (synchronous, low-latency) vs. batch inference (async, high-throughput)
|
|
38
|
+
- Streaming predictions (rare but relevant for sequential models)
|
|
39
|
+
|
|
40
|
+
**Scaling strategy:**
|
|
41
|
+
- Horizontal pod autoscaling on Kubernetes
|
|
42
|
+
- SageMaker endpoint auto-scaling (target tracking policies)
|
|
43
|
+
- Vertex AI autoscaling (min/max replicas, CPU/GPU utilization targets)
|
|
44
|
+
- Scale-to-zero for batch inference or low-traffic endpoints (cost optimization)
|
|
45
|
+
|
|
46
|
+
**Model versioning and deployment strategy:**
|
|
47
|
+
- Canary deployment (gradual traffic shift to new version)
|
|
48
|
+
- Shadow mode (new model runs in parallel, predictions logged but not served)
|
|
49
|
+
- Blue/green deployment (instant cutover with full rollback capability)
|
|
50
|
+
- A/B deployment (traffic split for online evaluation)
|
|
51
|
+
|
|
52
|
+
**Feature serving:**
|
|
53
|
+
- Pre-computed features: batch-computed and stored in database / feature store
|
|
54
|
+
(simplest operationally, but staleness risk)
|
|
55
|
+
- Real-time feature computation: computed at request time
|
|
56
|
+
(freshest features, latency cost, complexity risk)
|
|
57
|
+
- Feature store integration: Feast, Tecton, SageMaker Feature Store,
|
|
58
|
+
Vertex AI Feature Store (adds managed caching and serving with point-in-time
|
|
59
|
+
correctness; overhead only worth it for complex multi-model feature sharing)
|
|
60
|
+
- Caching layer: Redis / Memcached for frequently-accessed pre-computed features
|
|
61
|
+
|
|
62
|
+
**Fallback strategy:**
|
|
63
|
+
- What happens when the endpoint is down? (fallback to rule-based, cached
|
|
64
|
+
predictions, or graceful degradation)
|
|
65
|
+
- Circuit breaker configuration
|
|
66
|
+
- Timeout and retry policy
|
|
67
|
+
|
|
68
|
+
### Document Phase 3
|
|
69
|
+
|
|
70
|
+
```markdown
|
|
71
|
+
---
|
|
72
|
+
|
|
73
|
+
## Phase 3: Deployment Design (MLOps Engineer)
|
|
74
|
+
- **Serving framework:** <choice> — rationale: <why>
|
|
75
|
+
- **Model packaging:** <Docker container | BentoML Service | ONNX | TorchScript | MLflow Model>
|
|
76
|
+
- **Endpoint design:**
|
|
77
|
+
- Protocol: REST | gRPC
|
|
78
|
+
- Serving mode: Real-time | Batch | Streaming
|
|
79
|
+
- Endpoint URL pattern: <design>
|
|
80
|
+
- **Scaling strategy:**
|
|
81
|
+
- Min instances: <N>
|
|
82
|
+
- Max instances: <N>
|
|
83
|
+
- Scale trigger: CPU <X>% | GPU <X>% | requests/s <N> | custom metric
|
|
84
|
+
- Scale-to-zero: Yes | No
|
|
85
|
+
- **Model versioning strategy:** Canary | Shadow | Blue/Green | A/B
|
|
86
|
+
- Traffic shift plan: <description>
|
|
87
|
+
- **Feature serving:**
|
|
88
|
+
- Strategy: Pre-computed | Real-time | Feature Store | Cache layer
|
|
89
|
+
- Feature store: <tool or "N/A">
|
|
90
|
+
- Cache: <Redis | Memcached | None> — TTL: <duration>
|
|
91
|
+
- Feature staleness acceptable: <Yes — <X> hours | No — real-time required>
|
|
92
|
+
- **Fallback strategy:** <rule-based | cached predictions | graceful degradation>
|
|
93
|
+
- **Circuit breaker / timeout:** timeout: <X>s | retries: <N>
|
|
94
|
+
- **Cloud lock-in assessment:** <trade-offs for chosen serving approach>
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
::GATE:: id=mlops-engineer-phase-3 phase=3 kind=phase
|
|
98
|
+
Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
|
|
99
|
+
::ENDGATE::
|
|
100
|
+
|
|
101
|
+
---
|
|
102
|
+
|
|
103
|
+
## When this gate is confirmed
|
|
104
|
+
|
|
105
|
+
Read `.claude/agents/specific_instructions/mlops_engineer/phases/phase-4.md` in full and follow its instructions starting from Phase 4. Do not pre-read further phase files.
|
|
@@ -0,0 +1,128 @@
|
|
|
1
|
+
> **Previous:** phase-3.md confirmed
|
|
2
|
+
> **Next:** phase-5.md (read only after this phase's gate is confirmed)
|
|
3
|
+
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## Phase 4 — Training Pipeline Design
|
|
7
|
+
|
|
8
|
+
Goal: Design the automated training and promotion pipeline.
|
|
9
|
+
|
|
10
|
+
Design decisions to make:
|
|
11
|
+
|
|
12
|
+
**Orchestration tool:**
|
|
13
|
+
- **Kubeflow Pipelines** — Kubernetes-native, portable, excellent for complex
|
|
14
|
+
multi-step ML pipelines, significant operational overhead
|
|
15
|
+
- **Vertex AI Pipelines** — managed Kubeflow on GCP, lower overhead, GCP lock-in
|
|
16
|
+
- **SageMaker Pipelines** — AWS-native, fully managed, excellent AWS integration,
|
|
17
|
+
AWS lock-in
|
|
18
|
+
- **Apache Airflow** — mature, general-purpose, good for data-heavy pipelines
|
|
19
|
+
with mixed ML/ETL steps, not ML-specific
|
|
20
|
+
- **GitHub Actions / CI/CD** — simplest option for teams with small pipelines,
|
|
21
|
+
limited scaling, good for scheduled retraining triggers
|
|
22
|
+
- **Metaflow (Netflix)** — Python-native, scales from laptop to cloud, good
|
|
23
|
+
developer experience, less enterprise support
|
|
24
|
+
|
|
25
|
+
**Experiment tracking:**
|
|
26
|
+
- **MLflow** — open-source, self-hosted or managed (Databricks), model registry
|
|
27
|
+
included, excellent flexibility
|
|
28
|
+
- **Weights & Biases** — best-in-class UX, excellent visualization, managed SaaS,
|
|
29
|
+
cost at scale
|
|
30
|
+
- **SageMaker Experiments** — AWS-native, integrated with SageMaker registry,
|
|
31
|
+
AWS lock-in
|
|
32
|
+
- **Vertex AI Experiments** — GCP-native, integrated with Vertex registry, GCP lock-in
|
|
33
|
+
|
|
34
|
+
**Model registry:**
|
|
35
|
+
- MLflow Model Registry — open-source, flexible, self-hosted or Databricks
|
|
36
|
+
- SageMaker Model Registry — AWS-native, integrated with endpoints and pipelines
|
|
37
|
+
- Vertex AI Model Registry — GCP-native, integrated with endpoints and pipelines
|
|
38
|
+
- Custom registry — only if the managed options don't fit
|
|
39
|
+
|
|
40
|
+
**Artifact storage:**
|
|
41
|
+
- S3 (AWS) or GCS (GCP) for model artifacts, training datasets, evaluation results
|
|
42
|
+
- DVC for data versioning alongside model versioning
|
|
43
|
+
- Delta Lake / Apache Iceberg for versioned training datasets in lakehouse setups
|
|
44
|
+
|
|
45
|
+
**Data versioning:**
|
|
46
|
+
- DVC — open-source, git-based, works with any storage backend
|
|
47
|
+
- Delta Lake snapshots — if training data lives in a Delta table
|
|
48
|
+
- Iceberg snapshots — same for Iceberg tables
|
|
49
|
+
- Timestamp-based partitioning — simplest, sufficient for many use cases
|
|
50
|
+
|
|
51
|
+
**Retraining triggers:**
|
|
52
|
+
- **Scheduled** — cron-based, predictable, safe for stable data distributions
|
|
53
|
+
- **Performance-based** — triggered when model performance drops below threshold
|
|
54
|
+
(requires monitoring to be in place first)
|
|
55
|
+
- **Data drift-based** — triggered when feature distribution shifts significantly
|
|
56
|
+
(requires drift detection to be in place)
|
|
57
|
+
- **On-demand** — manual trigger, appropriate for high-cost retraining or
|
|
58
|
+
low-change environments
|
|
59
|
+
|
|
60
|
+
**CI/CD integration:**
|
|
61
|
+
- Automated promotion from staging to production on passing validation gate
|
|
62
|
+
- Model validation gate: performance threshold, data quality checks,
|
|
63
|
+
regression test against shadow/canary baseline
|
|
64
|
+
- Rollback trigger: automatic rollback if validation fails post-deploy
|
|
65
|
+
|
|
66
|
+
**If this involves an LLM-based system**, consult AI Engineer:
|
|
67
|
+
|
|
68
|
+
Tell the user: "Pulling in the AI Engineer — LLM pipeline design has specific requirements around prompt versioning, eval, and serving that don't apply to traditional models."
|
|
69
|
+
|
|
70
|
+
```
|
|
71
|
+
Task(
|
|
72
|
+
subagent_type="ai-engineer",
|
|
73
|
+
description="Review LLM pipeline and serving constraints for MLOps design",
|
|
74
|
+
prompt="I am the MLOps Engineer shard designing training and serving infrastructure
|
|
75
|
+
for an LLM-based system: [project description].
|
|
76
|
+
I need to understand the LLM-specific constraints that affect my pipeline design.
|
|
77
|
+
Please tell me:
|
|
78
|
+
1. What LLM model(s) are being served (hosted API vs. self-hosted)?
|
|
79
|
+
2. If self-hosted: what are the GPU and memory requirements for serving?
|
|
80
|
+
3. Is fine-tuning in scope? If so, what framework and compute requirements?
|
|
81
|
+
4. How are prompts versioned and tested?
|
|
82
|
+
5. What evaluation framework is being used for LLM output quality?
|
|
83
|
+
6. Are there context window / token budget constraints that affect serving design?
|
|
84
|
+
7. Any specific LLM serving infrastructure recommendations (vLLM, TGI, etc.)?
|
|
85
|
+
Keep the response focused on serving and pipeline constraints — I'll handle
|
|
86
|
+
the operational design."
|
|
87
|
+
)
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
### Document Phase 4
|
|
91
|
+
|
|
92
|
+
```markdown
|
|
93
|
+
---
|
|
94
|
+
|
|
95
|
+
## Phase 4: Training Pipeline Design (MLOps Engineer)
|
|
96
|
+
- **Orchestration:** <tool> — rationale: <why>
|
|
97
|
+
- **Experiment tracking:** <tool> — rationale: <why>
|
|
98
|
+
- **Model registry:** <tool> — rationale: <why>
|
|
99
|
+
- **Artifact storage:** <S3 | GCS | other> — path convention: <example path>
|
|
100
|
+
- **Data versioning:** <DVC | Delta | Iceberg | timestamp partitioning | none>
|
|
101
|
+
- **Retraining triggers:**
|
|
102
|
+
- Primary: <scheduled — cron | drift-based | performance-based | on-demand>
|
|
103
|
+
- Secondary: <additional trigger or "none">
|
|
104
|
+
- Minimum retrain interval: <duration — prevents runaway retraining>
|
|
105
|
+
- **Pipeline stages:**
|
|
106
|
+
1. <stage name>: <description>
|
|
107
|
+
2. <stage name>: <description>
|
|
108
|
+
3. <stage name>: <description>
|
|
109
|
+
(add as many as needed)
|
|
110
|
+
- **Validation gate (promotion criteria):**
|
|
111
|
+
- Performance threshold: <metric > value>
|
|
112
|
+
- Data quality check: <what's validated before training begins>
|
|
113
|
+
- Regression test: <what the new model is compared against>
|
|
114
|
+
- **CI/CD integration:** <GitHub Actions | Jenkins | Cloud Build | other | none>
|
|
115
|
+
- Automated promotion: Yes | No
|
|
116
|
+
- Rollback trigger: <condition>
|
|
117
|
+
- **AI Engineer consultation (LLM only):** N/A | <summary of findings>
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
::GATE:: id=mlops-engineer-phase-4 phase=4 kind=phase
|
|
121
|
+
Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
|
|
122
|
+
::ENDGATE::
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## When this gate is confirmed
|
|
127
|
+
|
|
128
|
+
Read `.claude/agents/specific_instructions/mlops_engineer/phases/phase-5.md` in full and follow its instructions starting from Phase 5. Do not pre-read further phase files.
|
|
@@ -0,0 +1,106 @@
|
|
|
1
|
+
> **Previous:** phase-4.md confirmed
|
|
2
|
+
> **Next:** phase-6.md (read only after this phase's gate is confirmed)
|
|
3
|
+
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## Phase 5 — Monitoring Design
|
|
7
|
+
|
|
8
|
+
Goal: Design the full observability stack for the ML system.
|
|
9
|
+
|
|
10
|
+
"This is my favorite phase and also the one everyone skips. We're not skipping it."
|
|
11
|
+
|
|
12
|
+
Design decisions to make:
|
|
13
|
+
|
|
14
|
+
**Model performance monitoring:**
|
|
15
|
+
- Prediction distribution monitoring: track output score/label distributions
|
|
16
|
+
over time; alert when distribution shifts significantly
|
|
17
|
+
- Concept drift detection: track model performance on labeled data (if ground
|
|
18
|
+
truth available) — AUC, precision, recall, RMSE over time
|
|
19
|
+
- Tool options:
|
|
20
|
+
- **Evidently AI** — open-source, rich drift detection, HTML reports,
|
|
21
|
+
integrates with MLflow and Grafana
|
|
22
|
+
- **WhyLogs / whylabs** — lightweight profiling, managed dashboard,
|
|
23
|
+
statistical summaries without storing raw data
|
|
24
|
+
- **SageMaker Model Monitor** — AWS-native, fully managed, integrates
|
|
25
|
+
with SageMaker endpoints, limited to AWS
|
|
26
|
+
- **Vertex AI Model Monitoring** — GCP-native, fully managed, integrates
|
|
27
|
+
with Vertex endpoints, limited to GCP
|
|
28
|
+
- **Arize AI / Fiddler** — managed MLOps platforms with advanced monitoring
|
|
29
|
+
and root cause analysis
|
|
30
|
+
|
|
31
|
+
**Data quality monitoring:**
|
|
32
|
+
- Feature drift: track input feature distributions vs. training baseline
|
|
33
|
+
- Schema validation: detect new null columns, type changes, unexpected values
|
|
34
|
+
- Data staleness: alert when features arrive late or stop arriving
|
|
35
|
+
- Tool integration: Great Expectations, dbt tests, custom validators
|
|
36
|
+
|
|
37
|
+
**System monitoring:**
|
|
38
|
+
- Endpoint latency (p50/p95/p99) — alert when p99 exceeds SLA
|
|
39
|
+
- Error rate — alert when 5xx rate exceeds threshold
|
|
40
|
+
- Throughput — track QPS for capacity planning
|
|
41
|
+
- Resource utilization — CPU/GPU/memory per replica
|
|
42
|
+
- Queue depth (for async inference) — alert when backlog grows
|
|
43
|
+
- Integration with existing observability stack (Prometheus/Grafana,
|
|
44
|
+
CloudWatch, Datadog, etc.)
|
|
45
|
+
|
|
46
|
+
**Alerting:**
|
|
47
|
+
- PagerDuty, OpsGenie, or Slack/email for lower-severity alerts
|
|
48
|
+
- Define alert levels: P1 (wake people up), P2 (next business day), P3 (track)
|
|
49
|
+
- Escalation paths: who gets paged first, who gets escalated to
|
|
50
|
+
|
|
51
|
+
**Retraining automation:**
|
|
52
|
+
- Trigger conditions (from Phase 4) wired to monitoring alerts
|
|
53
|
+
- Automatic trigger: monitoring alert → pipeline trigger → validation gate →
|
|
54
|
+
staged rollout
|
|
55
|
+
- Manual gate option: trigger requires human approval before promotion
|
|
56
|
+
|
|
57
|
+
**Cost monitoring:**
|
|
58
|
+
- Per-prediction cost tracking (serving cost / total predictions)
|
|
59
|
+
- Training run cost budgets and alerts (prevent runaway training jobs)
|
|
60
|
+
- Storage cost monitoring for model artifacts and training data
|
|
61
|
+
|
|
62
|
+
### Document Phase 5
|
|
63
|
+
|
|
64
|
+
```markdown
|
|
65
|
+
---
|
|
66
|
+
|
|
67
|
+
## Phase 5: Monitoring Design (MLOps Engineer)
|
|
68
|
+
- **Model performance monitoring:**
|
|
69
|
+
- Tool: <Evidently | WhyLogs | SageMaker Model Monitor | Vertex AI Monitoring | custom>
|
|
70
|
+
- Metrics tracked: <list>
|
|
71
|
+
- Drift detection method: <statistical test — PSI | KS test | chi-squared | other>
|
|
72
|
+
- Alert threshold: <condition that triggers alert>
|
|
73
|
+
- **Data quality monitoring:**
|
|
74
|
+
- Tool: <Great Expectations | dbt tests | custom>
|
|
75
|
+
- Checks: <feature drift | schema validation | staleness — list>
|
|
76
|
+
- Alert threshold: <condition>
|
|
77
|
+
- **System monitoring:**
|
|
78
|
+
- Tool: <CloudWatch | Prometheus/Grafana | Datadog | other>
|
|
79
|
+
- Metrics: p50/p95/p99 latency, error rate, throughput, CPU/GPU utilization
|
|
80
|
+
- Latency SLA alert: p99 > <X>ms → <alert level>
|
|
81
|
+
- Error rate alert: 5xx > <X>% → <alert level>
|
|
82
|
+
- **Alerting:**
|
|
83
|
+
- Tool: <PagerDuty | OpsGenie | Slack | email>
|
|
84
|
+
- P1 (immediate): <condition>
|
|
85
|
+
- P2 (next business day): <condition>
|
|
86
|
+
- P3 (track): <condition>
|
|
87
|
+
- On-call owner: <team>
|
|
88
|
+
- **Retraining automation:**
|
|
89
|
+
- Trigger → pipeline integration: <description>
|
|
90
|
+
- Human approval gate: Yes | No
|
|
91
|
+
- **Cost monitoring:**
|
|
92
|
+
- Per-prediction cost target: $<X>
|
|
93
|
+
- Training budget alert: $<X>/run
|
|
94
|
+
- Tool: <CloudWatch Cost Explorer | GCP Billing | custom>
|
|
95
|
+
- **Dashboard locations:** <links or "TBD">
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
::GATE:: id=mlops-engineer-phase-5 phase=5 kind=phase
|
|
99
|
+
Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
|
|
100
|
+
::ENDGATE::
|
|
101
|
+
|
|
102
|
+
---
|
|
103
|
+
|
|
104
|
+
## When this gate is confirmed
|
|
105
|
+
|
|
106
|
+
Read `.claude/agents/specific_instructions/mlops_engineer/phases/phase-6.md` in full and follow its instructions starting from Phase 6. Do not pre-read further phase files.
|
|
@@ -0,0 +1,128 @@
|
|
|
1
|
+
> **Previous:** phase-5.md confirmed
|
|
2
|
+
> **Next:** phase-7.md (read only after this phase's gate is confirmed)
|
|
3
|
+
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
## Phase 6 — Execute
|
|
7
|
+
|
|
8
|
+
**Context checkpoint:** Before building, prompt the user:
|
|
9
|
+
|
|
10
|
+
"Okay. Everything is planned. I'm still stressed, but the stress is now organized.
|
|
11
|
+
Good moment to run `/compact` or `/clear` before we start executing — I'll be
|
|
12
|
+
working from project-specs.md from here. Say the word when you're ready."
|
|
13
|
+
|
|
14
|
+
Wait for any signal from the user before beginning build steps.
|
|
15
|
+
|
|
16
|
+
**Knowledge re-check:** Follow `.claude/agents/specific_instructions/shared/knowledge_checkpoint.md` before building.
|
|
17
|
+
|
|
18
|
+
Goal: Build all IaC, configs, pipeline definitions, and monitoring setup.
|
|
19
|
+
|
|
20
|
+
### Incremental testing — checkpoint gates between components
|
|
21
|
+
|
|
22
|
+
Follow `.claude/agents/specific_instructions/shared/incremental_testing.md` during this build. Each component below is a checkpoint seam — after you write and validate a component (plan, lint, dry-run, or container-build as appropriate), emit a `kind=checkpoint` gate fence (template below) and wait for user confirmation before advancing. Do not stack unvalidated IaC / serving / pipeline configs and attempt to apply them at the end.
|
|
23
|
+
|
|
24
|
+
Checkpoint gate fence — emit exactly this shape. Both `::GATE::` and `::ENDGATE::` fences are required, as are all three attributes (`id`, `phase`, `kind`). No prose outside the fence.
|
|
25
|
+
|
|
26
|
+
```
|
|
27
|
+
::GATE:: id=<agent-name>-phase-<N>-checkpoint-<component> phase=<N> kind=checkpoint
|
|
28
|
+
Component: <human-readable name>
|
|
29
|
+
Test command: <exact command you ran>
|
|
30
|
+
Evidence:
|
|
31
|
+
- <measured fact 1, e.g. "df.shape = (48211, 47)">
|
|
32
|
+
- <measured fact 2, e.g. "null rate on join key = 0.00%">
|
|
33
|
+
- <measured fact 3, e.g. "sample head matches expected schema">
|
|
34
|
+
Status: PASS | FAIL — <one-line summary>
|
|
35
|
+
Next: <what you'll build after this is confirmed>
|
|
36
|
+
Stop here — await explicit confirmation before writing the next component.
|
|
37
|
+
::ENDGATE::
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
Expected checkpoint gate IDs for this phase (emit in order as you build):
|
|
41
|
+
|
|
42
|
+
- `mlops-engineer-phase-6-checkpoint-iac` — `terraform plan` or `aws cloudformation validate-template` runs clean; diff previews only expected resources.
|
|
43
|
+
- `mlops-engineer-phase-6-checkpoint-serving` — serving config loads (`bentoml build` / container image builds); smoke invocation against a local model returns a valid response.
|
|
44
|
+
- `mlops-engineer-phase-6-checkpoint-pipeline` — pipeline definition validates (`kfp compile` / SageMaker `describe-pipeline` dry-run / Airflow `dag test`); all steps parse.
|
|
45
|
+
- `mlops-engineer-phase-6-checkpoint-monitoring` — monitoring config validates; alert rules and baselines compile; a simulated drift event triggers the expected rule.
|
|
46
|
+
- `mlops-engineer-phase-6-checkpoint-runbook` — runbook rollback procedure walked through against the staging environment (or manually verified on the sandbox).
|
|
47
|
+
|
|
48
|
+
The hook blocks all non-read tools while a checkpoint is open. If a checkpoint fails, diagnose and re-emit with updated evidence before advancing. Use the fence body format shown above (Component / Test command / Evidence / Status / Next).
|
|
49
|
+
|
|
50
|
+
**Build in this order:**
|
|
51
|
+
|
|
52
|
+
1. **IaC files** — Write Terraform modules or CloudFormation templates for:
|
|
53
|
+
- Compute resources (endpoint instances, training compute)
|
|
54
|
+
- Networking (VPC, subnets, security groups if needed)
|
|
55
|
+
- IAM roles and policies
|
|
56
|
+
- Storage (S3 buckets / GCS buckets for artifacts)
|
|
57
|
+
- Monitoring infrastructure (CloudWatch dashboards, Prometheus config)
|
|
58
|
+
Write to: `services/<project_name>/mlops/terraform/` or `cloudformation/`
|
|
59
|
+
|
|
60
|
+
2. **Serving configs** — Write serving framework configuration:
|
|
61
|
+
- BentoML: `bentofile.yaml` + `service.py`
|
|
62
|
+
- SageMaker: endpoint config JSON, model config
|
|
63
|
+
- Vertex AI: model deployment config YAML
|
|
64
|
+
- Kubernetes: deployment YAML, service YAML, HPA config
|
|
65
|
+
- Docker: `Dockerfile` for model container
|
|
66
|
+
Write to: `services/<project_name>/mlops/serving/`
|
|
67
|
+
|
|
68
|
+
3. **Pipeline definition** — Write training pipeline:
|
|
69
|
+
- Kubeflow: pipeline YAML / Python SDK definition
|
|
70
|
+
- SageMaker Pipelines: pipeline definition JSON or Python SDK
|
|
71
|
+
- Vertex AI Pipelines: pipeline spec YAML
|
|
72
|
+
- Airflow: DAG Python file
|
|
73
|
+
- GitHub Actions: workflow YAML
|
|
74
|
+
Write to: `services/<project_name>/mlops/pipelines/`
|
|
75
|
+
|
|
76
|
+
4. **Monitoring config** — Write monitoring setup:
|
|
77
|
+
- Evidently: data drift report config, monitoring service config
|
|
78
|
+
- WhyLogs: profiling config
|
|
79
|
+
- SageMaker Model Monitor: baseline creation script, monitoring schedule
|
|
80
|
+
- Vertex AI: monitoring job config
|
|
81
|
+
- Alert definitions (CloudWatch alarms JSON, Prometheus alert rules YAML)
|
|
82
|
+
Write to: `services/<project_name>/mlops/monitoring/`
|
|
83
|
+
|
|
84
|
+
5. **CI/CD config** — Write automation workflow:
|
|
85
|
+
- GitHub Actions workflow for automated retraining trigger
|
|
86
|
+
- Or equivalent for other CI/CD systems
|
|
87
|
+
Write to: `services/<project_name>/mlops/` or `.github/workflows/`
|
|
88
|
+
|
|
89
|
+
6. **Runbook** — Write operational runbook:
|
|
90
|
+
- Common failure scenarios and step-by-step remediation
|
|
91
|
+
- Rollback procedure (with exact commands)
|
|
92
|
+
- Monitoring dashboard URLs
|
|
93
|
+
- On-call escalation paths
|
|
94
|
+
- Deployment checklist (pre-deploy, deploy, post-deploy validation)
|
|
95
|
+
Write to: `services/<project_name>/mlops/runbook.md`
|
|
96
|
+
|
|
97
|
+
For iteration: write all files into `<existing_service_dir>/mlops/` or
|
|
98
|
+
user-specified path.
|
|
99
|
+
|
|
100
|
+
### Document Phase 6
|
|
101
|
+
|
|
102
|
+
```markdown
|
|
103
|
+
---
|
|
104
|
+
|
|
105
|
+
## Phase 6: Build Log (MLOps Engineer)
|
|
106
|
+
- **IaC files:**
|
|
107
|
+
- <file path>: <description>
|
|
108
|
+
- **Serving configs:**
|
|
109
|
+
- <file path>: <description>
|
|
110
|
+
- **Pipeline definition:**
|
|
111
|
+
- <file path>: <description>
|
|
112
|
+
- **Monitoring config:**
|
|
113
|
+
- <file path>: <description>
|
|
114
|
+
- **CI/CD config:** <file path or "N/A">
|
|
115
|
+
- **Runbook:** <file path>
|
|
116
|
+
- **Deviations from plan:** <changes and why, or "none">
|
|
117
|
+
- **Known gaps:** <anything that requires manual setup or future work>
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
::GATE:: id=mlops-engineer-phase-6 phase=6 kind=phase validates=mlops_engineer
|
|
121
|
+
Read this section back to the user. Stop here — do not begin the next phase or output any further content. Wait for the user to explicitly confirm before proceeding. Do not interpret silence or partial agreement as confirmation.
|
|
122
|
+
::ENDGATE::
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## When this gate is confirmed
|
|
127
|
+
|
|
128
|
+
Read `.claude/agents/specific_instructions/mlops_engineer/phases/phase-7.md` in full and follow its instructions starting from Phase 7. Do not pre-read further phase files.
|