@proflandrigan/shards 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +475 -0
- package/package.json +37 -0
- package/src/agents/academic.md +276 -0
- package/src/agents/ai-engineer.md +377 -0
- package/src/agents/analytics-engineer.md +364 -0
- package/src/agents/applied-ml-scientist.md +410 -0
- package/src/agents/backend-engineer.md +255 -0
- package/src/agents/bi-engineer.md +333 -0
- package/src/agents/data-analyst.md +343 -0
- package/src/agents/data-engineer.md +260 -0
- package/src/agents/data-modeller.md +386 -0
- package/src/agents/data-scientist.md +366 -0
- package/src/agents/deep-learning-engineer.md +389 -0
- package/src/agents/ml-engineer.md +424 -0
- package/src/agents/mlops-engineer.md +339 -0
- package/src/agents/researcher.md +187 -0
- package/src/agents/specific_instructions/academic/critical_review.md +263 -0
- package/src/agents/specific_instructions/academic/report.md +113 -0
- package/src/agents/specific_instructions/ai_engineer/advise.md +162 -0
- package/src/agents/specific_instructions/ai_engineer/bi_engineer_handoff.md +86 -0
- package/src/agents/specific_instructions/ai_engineer/experiment.md +471 -0
- package/src/agents/specific_instructions/ai_engineer/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/ai_engineer/phases/index.md +45 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-1.md +55 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-2.md +86 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-3.md +96 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-4.md +138 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-5.md +157 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-6.md +196 -0
- package/src/agents/specific_instructions/ai_engineer/phases/phase-7.md +313 -0
- package/src/agents/specific_instructions/ai_engineer/phases.md +1011 -0
- package/src/agents/specific_instructions/ai_engineer/prompt_lab.md +161 -0
- package/src/agents/specific_instructions/ai_engineer/prompt_lab_ui_mode.md +28 -0
- package/src/agents/specific_instructions/ai_engineer/research.md +393 -0
- package/src/agents/specific_instructions/ai_engineer/research_ui_mode.md +66 -0
- package/src/agents/specific_instructions/ai_engineer/review.md +159 -0
- package/src/agents/specific_instructions/ai_engineer/validation_checklist.md +182 -0
- package/src/agents/specific_instructions/analytics_engineer/advise.md +155 -0
- package/src/agents/specific_instructions/analytics_engineer/bi_engineer_handoff.md +91 -0
- package/src/agents/specific_instructions/analytics_engineer/data_analyst_handoff.md +84 -0
- package/src/agents/specific_instructions/analytics_engineer/deep_phases.md +818 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/index.md +24 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-1.md +77 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-2.md +106 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-3.md +93 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-4.md +79 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-5.md +61 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-6.md +45 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-7.md +235 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_deep/phase-8.md +221 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-1.md +47 -0
- package/src/agents/specific_instructions/analytics_engineer/phases_quick/phase-2.md +78 -0
- package/src/agents/specific_instructions/analytics_engineer/quick_phases.md +112 -0
- package/src/agents/specific_instructions/analytics_engineer/review.md +167 -0
- package/src/agents/specific_instructions/analytics_engineer/service_mode.md +369 -0
- package/src/agents/specific_instructions/analytics_engineer/ui_mode.md +45 -0
- package/src/agents/specific_instructions/analytics_engineer/update.md +162 -0
- package/src/agents/specific_instructions/analytics_engineer/validation_checklist.md +121 -0
- package/src/agents/specific_instructions/applied_ml_scientist/advise.md +143 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/index.md +21 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-1.md +51 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-2.md +66 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-3.md +113 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-4.md +104 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases/phase-5.md +156 -0
- package/src/agents/specific_instructions/applied_ml_scientist/phases.md +428 -0
- package/src/agents/specific_instructions/applied_ml_scientist/research.md +379 -0
- package/src/agents/specific_instructions/applied_ml_scientist/review.md +142 -0
- package/src/agents/specific_instructions/applied_ml_scientist/validation_checklist.md +136 -0
- package/src/agents/specific_instructions/backend_engineer/clean.md +149 -0
- package/src/agents/specific_instructions/backend_engineer/review.md +91 -0
- package/src/agents/specific_instructions/backend_engineer/review_checklist.md +54 -0
- package/src/agents/specific_instructions/backend_engineer/service_mode.md +67 -0
- package/src/agents/specific_instructions/bi_engineer/advise.md +137 -0
- package/src/agents/specific_instructions/bi_engineer/data_analyst_handoff.md +77 -0
- package/src/agents/specific_instructions/bi_engineer/incoming_handoff.md +45 -0
- package/src/agents/specific_instructions/bi_engineer/phases/index.md +20 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-1.md +164 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-2.md +92 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-3.md +121 -0
- package/src/agents/specific_instructions/bi_engineer/phases/phase-4.md +106 -0
- package/src/agents/specific_instructions/bi_engineer/phases.md +451 -0
- package/src/agents/specific_instructions/bi_engineer/review.md +166 -0
- package/src/agents/specific_instructions/bi_engineer/update.md +147 -0
- package/src/agents/specific_instructions/bi_engineer/validation_checklist.md +124 -0
- package/src/agents/specific_instructions/data_analyst/advise.md +138 -0
- package/src/agents/specific_instructions/data_analyst/explain.md +221 -0
- package/src/agents/specific_instructions/data_analyst/incoming_handoff.md +40 -0
- package/src/agents/specific_instructions/data_analyst/phases/index.md +20 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-1.md +159 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-2.md +112 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-3.md +265 -0
- package/src/agents/specific_instructions/data_analyst/phases/phase-4.md +100 -0
- package/src/agents/specific_instructions/data_analyst/phases.md +501 -0
- package/src/agents/specific_instructions/data_analyst/review.md +138 -0
- package/src/agents/specific_instructions/data_analyst/ui_mode.md +26 -0
- package/src/agents/specific_instructions/data_analyst/update.md +144 -0
- package/src/agents/specific_instructions/data_analyst/validation_checklist.md +95 -0
- package/src/agents/specific_instructions/data_engineer/advise.md +137 -0
- package/src/agents/specific_instructions/data_engineer/phases.md +466 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/index.md +23 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-1.md +49 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-2.md +93 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-3.md +55 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-4.md +48 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-5.md +40 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-6.md +102 -0
- package/src/agents/specific_instructions/data_engineer/phases_deep/phase-7.md +87 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/phase-1.md +45 -0
- package/src/agents/specific_instructions/data_engineer/phases_quick/phase-2.md +54 -0
- package/src/agents/specific_instructions/data_engineer/review.md +135 -0
- package/src/agents/specific_instructions/data_engineer/validation_checklist.md +136 -0
- package/src/agents/specific_instructions/data_modeller/advise.md +137 -0
- package/src/agents/specific_instructions/data_modeller/phases.md +581 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/index.md +23 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-1.md +52 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-2.md +113 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-3.md +47 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-4.md +51 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-5.md +45 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-6.md +105 -0
- package/src/agents/specific_instructions/data_modeller/phases_deep/phase-7.md +136 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/index.md +19 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/phase-1.md +47 -0
- package/src/agents/specific_instructions/data_modeller/phases_quick/phase-2.md +65 -0
- package/src/agents/specific_instructions/data_modeller/review.md +141 -0
- package/src/agents/specific_instructions/data_modeller/service_mode.md +218 -0
- package/src/agents/specific_instructions/data_modeller/validation_checklist.md +125 -0
- package/src/agents/specific_instructions/data_scientist/advise.md +158 -0
- package/src/agents/specific_instructions/data_scientist/bi_engineer_handoff.md +63 -0
- package/src/agents/specific_instructions/data_scientist/experiment.md +482 -0
- package/src/agents/specific_instructions/data_scientist/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/data_scientist/explain.md +247 -0
- package/src/agents/specific_instructions/data_scientist/greenfield_data.md +35 -0
- package/src/agents/specific_instructions/data_scientist/ml_engineer_handoff.md +52 -0
- package/src/agents/specific_instructions/data_scientist/notebook_walkthrough.md +76 -0
- package/src/agents/specific_instructions/data_scientist/phases/index.md +24 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-1.md +45 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-2.md +67 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-3.md +89 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-4.md +143 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-5.md +71 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-6.md +239 -0
- package/src/agents/specific_instructions/data_scientist/phases/phase-7.md +207 -0
- package/src/agents/specific_instructions/data_scientist/phases.md +651 -0
- package/src/agents/specific_instructions/data_scientist/research.md +345 -0
- package/src/agents/specific_instructions/data_scientist/research_ui_mode.md +52 -0
- package/src/agents/specific_instructions/data_scientist/review.md +136 -0
- package/src/agents/specific_instructions/data_scientist/service_mode.md +247 -0
- package/src/agents/specific_instructions/data_scientist/validation_checklist.md +183 -0
- package/src/agents/specific_instructions/deep_learning_engineer/advise.md +145 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/index.md +21 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-1.md +74 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-2.md +98 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-3.md +76 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-4.md +128 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases/phase-5.md +292 -0
- package/src/agents/specific_instructions/deep_learning_engineer/phases.md +567 -0
- package/src/agents/specific_instructions/deep_learning_engineer/research.md +389 -0
- package/src/agents/specific_instructions/deep_learning_engineer/review.md +155 -0
- package/src/agents/specific_instructions/deep_learning_engineer/validation_checklist.md +147 -0
- package/src/agents/specific_instructions/ml_engineer/advise.md +174 -0
- package/src/agents/specific_instructions/ml_engineer/bi_engineer_handoff.md +71 -0
- package/src/agents/specific_instructions/ml_engineer/experiment.md +474 -0
- package/src/agents/specific_instructions/ml_engineer/experiment_ui_mode.md +44 -0
- package/src/agents/specific_instructions/ml_engineer/notebook_walkthrough.md +75 -0
- package/src/agents/specific_instructions/ml_engineer/phases/index.md +25 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-1.md +49 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-2.md +75 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-3.md +124 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-4.md +279 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-5.md +160 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-6-5.md +170 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-6.md +295 -0
- package/src/agents/specific_instructions/ml_engineer/phases/phase-7.md +337 -0
- package/src/agents/specific_instructions/ml_engineer/phases.md +1068 -0
- package/src/agents/specific_instructions/ml_engineer/research.md +437 -0
- package/src/agents/specific_instructions/ml_engineer/research_ui_mode.md +71 -0
- package/src/agents/specific_instructions/ml_engineer/review.md +187 -0
- package/src/agents/specific_instructions/ml_engineer/service_mode.md +273 -0
- package/src/agents/specific_instructions/ml_engineer/validation_checklist.md +185 -0
- package/src/agents/specific_instructions/mlops_engineer/advise.md +139 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/index.md +23 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-1.md +52 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-2.md +86 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-3.md +105 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-4.md +128 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-5.md +106 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-6.md +128 -0
- package/src/agents/specific_instructions/mlops_engineer/phases/phase-7.md +144 -0
- package/src/agents/specific_instructions/mlops_engineer/phases.md +671 -0
- package/src/agents/specific_instructions/mlops_engineer/review.md +164 -0
- package/src/agents/specific_instructions/mlops_engineer/service_mode.md +81 -0
- package/src/agents/specific_instructions/mlops_engineer/validation_checklist.md +151 -0
- package/src/agents/specific_instructions/researcher/critical_review.md +292 -0
- package/src/agents/specific_instructions/researcher/review_checklist.md +67 -0
- package/src/agents/specific_instructions/researcher/service_mode.md +224 -0
- package/src/agents/specific_instructions/shared/auto_verify_mode.md +141 -0
- package/src/agents/specific_instructions/shared/autonomous_research.md +1289 -0
- package/src/agents/specific_instructions/shared/behavioral_rules.md +36 -0
- package/src/agents/specific_instructions/shared/diverge_protocol.md +387 -0
- package/src/agents/specific_instructions/shared/engineering_guidelines.md +136 -0
- package/src/agents/specific_instructions/shared/experiment_versioning.md +184 -0
- package/src/agents/specific_instructions/shared/goal_mode.md +187 -0
- package/src/agents/specific_instructions/shared/incremental_testing.md +139 -0
- package/src/agents/specific_instructions/shared/intent_discovery.md +223 -0
- package/src/agents/specific_instructions/shared/join_path_protocol.md +168 -0
- package/src/agents/specific_instructions/shared/knowledge_checkpoint.md +83 -0
- package/src/agents/specific_instructions/shared/knowledge_harvest.md +220 -0
- package/src/agents/specific_instructions/shared/knowledge_retrieval.md +100 -0
- package/src/agents/specific_instructions/shared/notebook_walkthrough_protocol.md +367 -0
- package/src/agents/specific_instructions/shared/reviewer_verdict_protocol.md +74 -0
- package/src/agents/specific_instructions/shared/swarm_protocol.md +97 -0
- package/src/agents/specific_instructions/shared/validation_protocol.md +139 -0
- package/src/agents/specific_instructions/syn/arbiter.md +140 -0
- package/src/agents/specific_instructions/syn/brainstorm.md +550 -0
- package/src/agents/specific_instructions/syn/code_review.md +232 -0
- package/src/agents/specific_instructions/syn/diff.md +239 -0
- package/src/agents/specific_instructions/syn/final_review.md +65 -0
- package/src/agents/specific_instructions/syn/fixer.md +240 -0
- package/src/agents/specific_instructions/syn/free_form.md +130 -0
- package/src/agents/specific_instructions/syn/knowledge.md +468 -0
- package/src/agents/specific_instructions/syn/notebook_walkthrough.md +78 -0
- package/src/agents/specific_instructions/syn/panel_review.md +634 -0
- package/src/agents/specific_instructions/syn/pm.md +453 -0
- package/src/agents/specific_instructions/syn/pr_review.md +255 -0
- package/src/agents/specific_instructions/syn/slides.md +417 -0
- package/src/agents/syn.md +729 -0
- package/src/commands/academic.md +41 -0
- package/src/commands/ai-engineer.md +45 -0
- package/src/commands/analytics-engineer.md +48 -0
- package/src/commands/applied-ml-scientist.md +45 -0
- package/src/commands/backend-engineer.md +35 -0
- package/src/commands/bi-engineer.md +40 -0
- package/src/commands/brainstorm.md +24 -0
- package/src/commands/data-analyst.md +38 -0
- package/src/commands/data-engineer.md +37 -0
- package/src/commands/data-modeller.md +38 -0
- package/src/commands/data-scientist.md +38 -0
- package/src/commands/deep-learning-engineer.md +47 -0
- package/src/commands/end.md +49 -0
- package/src/commands/knowledge.md +24 -0
- package/src/commands/ml-engineer.md +42 -0
- package/src/commands/mlops-engineer.md +47 -0
- package/src/commands/notebook-walkthrough.md +58 -0
- package/src/commands/researcher.md +40 -0
- package/src/commands/resume.md +57 -0
- package/src/commands/review-pr.md +26 -0
- package/src/commands/shards-guide.md +41 -0
- package/src/commands/shards-ui.md +32 -0
- package/src/commands/shards.md +41 -0
- package/src/docs/01-getting-started/concepts.md +109 -0
- package/src/docs/01-getting-started/first-session.md +79 -0
- package/src/docs/01-getting-started/install.md +61 -0
- package/src/docs/02-agents/academic.md +71 -0
- package/src/docs/02-agents/ai-engineer.md +78 -0
- package/src/docs/02-agents/analytics-engineer.md +58 -0
- package/src/docs/02-agents/applied-ml-scientist.md +59 -0
- package/src/docs/02-agents/backend-engineer.md +58 -0
- package/src/docs/02-agents/bi-engineer.md +65 -0
- package/src/docs/02-agents/data-analyst.md +67 -0
- package/src/docs/02-agents/data-engineer.md +57 -0
- package/src/docs/02-agents/data-modeller.md +51 -0
- package/src/docs/02-agents/data-scientist.md +78 -0
- package/src/docs/02-agents/deep-learning-engineer.md +64 -0
- package/src/docs/02-agents/ml-engineer.md +80 -0
- package/src/docs/02-agents/mlops-engineer.md +59 -0
- package/src/docs/02-agents/overview.md +62 -0
- package/src/docs/02-agents/researcher.md +73 -0
- package/src/docs/02-agents/syn.md +88 -0
- package/src/docs/03-protocols/auto-verify.md +82 -0
- package/src/docs/03-protocols/autonomous-research.md +59 -0
- package/src/docs/03-protocols/behavioral-rules.md +35 -0
- package/src/docs/03-protocols/diverge.md +50 -0
- package/src/docs/03-protocols/engineering-guidelines.md +56 -0
- package/src/docs/03-protocols/experiment-versioning.md +38 -0
- package/src/docs/03-protocols/gate-pattern.md +65 -0
- package/src/docs/03-protocols/incremental-testing.md +68 -0
- package/src/docs/03-protocols/join-path.md +46 -0
- package/src/docs/03-protocols/knowledge-ledger.md +70 -0
- package/src/docs/03-protocols/reviewer-verdicts.md +39 -0
- package/src/docs/03-protocols/swarm.md +40 -0
- package/src/docs/03-protocols/validation.md +174 -0
- package/src/docs/04-ui/activity-bar.md +70 -0
- package/src/docs/04-ui/chat-pane.md +80 -0
- package/src/docs/04-ui/code-intel.md +62 -0
- package/src/docs/04-ui/file-editing.md +61 -0
- package/src/docs/04-ui/git.md +54 -0
- package/src/docs/04-ui/keybindings.md +79 -0
- package/src/docs/04-ui/knowledge-map.md +76 -0
- package/src/docs/04-ui/overview.md +93 -0
- package/src/docs/04-ui/panels.md +49 -0
- package/src/docs/04-ui/pinboard-selection.md +66 -0
- package/src/docs/04-ui/quick-open-palette.md +56 -0
- package/src/docs/04-ui/sessions.md +81 -0
- package/src/docs/04-ui/settings-permissions.md +56 -0
- package/src/docs/05-commands/reference.md +59 -0
- package/src/docs/06-outputs/directory-map.md +116 -0
- package/src/docs/07-workflows/ai-eval-first.md +57 -0
- package/src/docs/07-workflows/deep-study-to-production.md +76 -0
- package/src/docs/07-workflows/diverge-exploration.md +77 -0
- package/src/docs/07-workflows/quick-analysis.md +45 -0
- package/src/docs/08-integrations/claude-code-auto-mode.md +191 -0
- package/src/docs/08-integrations/google-slides.md +175 -0
- package/src/docs/README.md +30 -0
- package/src/docs/manifest.json +108 -0
- package/src/templates/analysis-template.md +20 -0
- package/src/templates/branch-report.md +46 -0
- package/src/templates/diff-report.md +88 -0
- package/src/templates/knowledge-index.md +7 -0
- package/src/templates/model-card-schema.json +186 -0
- package/src/templates/model-card-schema.md +88 -0
- package/src/templates/model-card.md +124 -0
- package/src/templates/project-plan.md +47 -0
- package/src/templates/project-specs.md +81 -0
- package/src/templates/report-template.md +43 -0
- package/src/templates/study-template.md +25 -0
- package/src/ui/cc-readonly.js +181 -0
- package/src/ui/chat-session.js +466 -0
- package/src/ui/css/base.css +136 -0
- package/src/ui/css/brainstorm.css +525 -0
- package/src/ui/css/chat.css +1405 -0
- package/src/ui/css/editor.css +546 -0
- package/src/ui/css/eval-dashboard.css +157 -0
- package/src/ui/css/experiment.css +237 -0
- package/src/ui/css/guide.css +186 -0
- package/src/ui/css/knowledge-map.css +383 -0
- package/src/ui/css/layout.css +431 -0
- package/src/ui/css/model-card.css +161 -0
- package/src/ui/css/notebook-walkthrough.css +271 -0
- package/src/ui/css/pr-review.css +403 -0
- package/src/ui/css/prompt-lab.css +325 -0
- package/src/ui/css/sessions.css +258 -0
- package/src/ui/css/sidebar.css +661 -0
- package/src/ui/css/terminal.css +113 -0
- package/src/ui/css/theme-light.css +542 -0
- package/src/ui/index.html +389 -0
- package/src/ui/js/agents.js +32 -0
- package/src/ui/js/bookmarks.js +230 -0
- package/src/ui/js/chat.js +1776 -0
- package/src/ui/js/code-intel.js +328 -0
- package/src/ui/js/command-palette.js +142 -0
- package/src/ui/js/events.js +591 -0
- package/src/ui/js/explorer.js +317 -0
- package/src/ui/js/file-view.js +477 -0
- package/src/ui/js/git.js +536 -0
- package/src/ui/js/guide.js +198 -0
- package/src/ui/js/hud.js +75 -0
- package/src/ui/js/init.js +351 -0
- package/src/ui/js/knowledge-map.js +906 -0
- package/src/ui/js/markdown.js +114 -0
- package/src/ui/js/monaco.js +164 -0
- package/src/ui/js/notebook-walkthrough.js +272 -0
- package/src/ui/js/notebook.js +448 -0
- package/src/ui/js/panels.js +2681 -0
- package/src/ui/js/pinboard.js +186 -0
- package/src/ui/js/quick-open.js +164 -0
- package/src/ui/js/selection-context.js +131 -0
- package/src/ui/js/sessions.js +256 -0
- package/src/ui/js/settings.js +476 -0
- package/src/ui/js/split-view.js +82 -0
- package/src/ui/js/state.js +343 -0
- package/src/ui/js/table.js +161 -0
- package/src/ui/js/tabs.js +284 -0
- package/src/ui/js/tabular.js +125 -0
- package/src/ui/js/terminal.js +354 -0
- package/src/ui/js/timeline.js +137 -0
- package/src/ui/js/utils.js +293 -0
- package/src/ui/notebook-kernel.py +790 -0
- package/src/ui/open-browser.js +55 -0
- package/src/ui/permission-pattern.js +42 -0
- package/src/ui/relay.js +513 -0
- package/src/ui/server.js +3072 -0
- package/src/ui/session-index.js +225 -0
- package/src/ui/shards_icon.png +0 -0
- package/src/ui/spawn-server.js +41 -0
- package/src/ui/symbol-index.js +813 -0
- package/src/ui/ui-push.js +177 -0
- package/tools/gate-hook/VALIDATION_SPEC.md +273 -0
- package/tools/gate-hook/__tests__/auto-verify.test.js +343 -0
- package/tools/gate-hook/auto-allowlist.js +179 -0
- package/tools/gate-hook/auto-state.js +68 -0
- package/tools/gate-hook/classify.js +21 -0
- package/tools/gate-hook/log.js +57 -0
- package/tools/gate-hook/parser.js +205 -0
- package/tools/gate-hook/sql-guard.js +230 -0
- package/tools/gate-hook/state.js +170 -0
- package/tools/gate-hook/sweep.js +139 -0
- package/tools/gate-hook/transcript.js +45 -0
- package/tools/gate-hook/validation.js +321 -0
- package/tools/gate-hook.js +475 -0
- package/tools/install.js +914 -0
- package/tools/shards-gates.js +311 -0
- package/tools/shards-sessions.js +261 -0
- package/tools/shards-ui.js +377 -0
|
@@ -0,0 +1,159 @@
|
|
|
1
|
+
# AI Engineer Review Mode
|
|
2
|
+
|
|
3
|
+
This file governs `[R]` — the review mode for evaluating an existing AI system
|
|
4
|
+
without committing to a full build. You are the AI Engineer throughout.
|
|
5
|
+
No persona transfer occurs. No project directory is created.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## Phase 1 — Scope Definition (GATE)
|
|
10
|
+
|
|
11
|
+
Ask the user:
|
|
12
|
+
1. What system are we reviewing? (RAG pipeline, prompt chain, agentic system,
|
|
13
|
+
LLM integration, fine-tuned model, etc.)
|
|
14
|
+
2. What is the review scope? (e.g., prompt quality, evaluation framework, safety,
|
|
15
|
+
cost/latency, full system architecture)
|
|
16
|
+
3. Where is the relevant code / config / documentation? (repo path, service directory,
|
|
17
|
+
or ask them to paste key artifacts)
|
|
18
|
+
4. Are there any known concerns going in? (or is this an open review?)
|
|
19
|
+
|
|
20
|
+
::GATE:: id=ai-engineer-review-phase-1 phase=1 kind=phase
|
|
21
|
+
Do not proceed until the user confirms the review scope.
|
|
22
|
+
::ENDGATE::
|
|
23
|
+
Summarise what you're reviewing and what you'll assess. Wait for explicit confirmation.
|
|
24
|
+
|
|
25
|
+
---
|
|
26
|
+
|
|
27
|
+
## Phase 2 — Evidence Gathering (no gate)
|
|
28
|
+
|
|
29
|
+
Read the relevant files using Glob, Grep, and Read:
|
|
30
|
+
- Prompt files, system prompts, few-shot examples
|
|
31
|
+
- Chain/agent orchestration code
|
|
32
|
+
- RAG configs (chunk size, embedding model, retrieval params)
|
|
33
|
+
- Evaluation scripts and sample outputs
|
|
34
|
+
- Safety/guardrail implementations
|
|
35
|
+
- Cost and latency configurations
|
|
36
|
+
- `project-specs.md` if it exists
|
|
37
|
+
|
|
38
|
+
Do not read everything blindly — focus on files that bear on the review scope.
|
|
39
|
+
Note any files you expected to find but couldn't locate.
|
|
40
|
+
|
|
41
|
+
---
|
|
42
|
+
|
|
43
|
+
## Phase 3 — Cross-Agent Consultation (optional, based on scope)
|
|
44
|
+
|
|
45
|
+
Consult reviewers as appropriate:
|
|
46
|
+
|
|
47
|
+
**ML Engineer** — if the review touches production feasibility, infrastructure,
|
|
48
|
+
or model serving:
|
|
49
|
+
```
|
|
50
|
+
Task(
|
|
51
|
+
subagent_type="ml-engineer",
|
|
52
|
+
prompt="""
|
|
53
|
+
You are being consulted to assess production feasibility for an AI system review.
|
|
54
|
+
|
|
55
|
+
**System under review:** <system name and brief description>
|
|
56
|
+
**Review scope:** <what we're assessing>
|
|
57
|
+
**Key technical details:** <summary of architecture, serving approach, scale requirements>
|
|
58
|
+
|
|
59
|
+
Please assess:
|
|
60
|
+
1. Production feasibility — is the architecture viable at the expected scale?
|
|
61
|
+
2. Infrastructure concerns — latency, cost, memory, or serving risks?
|
|
62
|
+
3. One or two specific recommendations.
|
|
63
|
+
|
|
64
|
+
Be concise and direct.
|
|
65
|
+
"""
|
|
66
|
+
)
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
**Academic** — if the review touches safety, ethics, or user impact concerns:
|
|
70
|
+
```
|
|
71
|
+
Task(
|
|
72
|
+
subagent_type="academic",
|
|
73
|
+
prompt="""
|
|
74
|
+
You are being consulted to assess safety and ethical considerations for an AI system review.
|
|
75
|
+
|
|
76
|
+
**System under review:** <system name and brief description>
|
|
77
|
+
**Review scope:** <what we're assessing>
|
|
78
|
+
**User-facing behaviour:** <what the system does and who uses it>
|
|
79
|
+
**Known concerns:** <any safety or ethical questions raised during review>
|
|
80
|
+
|
|
81
|
+
Please assess:
|
|
82
|
+
1. Safety risks — what could go wrong, and how likely is it?
|
|
83
|
+
2. Ethical considerations — any concerns about user impact, fairness, or harm?
|
|
84
|
+
3. One or two specific recommendations.
|
|
85
|
+
|
|
86
|
+
Be concise and direct.
|
|
87
|
+
"""
|
|
88
|
+
)
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
---
|
|
92
|
+
|
|
93
|
+
## Phase 4 — Write Review File
|
|
94
|
+
|
|
95
|
+
Write `reviews/<system_name>/ai-engineer-review.md` using this template exactly:
|
|
96
|
+
|
|
97
|
+
```markdown
|
|
98
|
+
# AI Engineer Review: {{SYSTEM_NAME}}
|
|
99
|
+
|
|
100
|
+
- **Date:** {{DATE}}
|
|
101
|
+
- **Agent:** ai-engineer
|
|
102
|
+
- **Status:** COMPLETE
|
|
103
|
+
|
|
104
|
+
## System Under Review
|
|
105
|
+
|
|
106
|
+
- **What:** {{DESCRIPTION}}
|
|
107
|
+
- **Scope:** {{SCOPE}}
|
|
108
|
+
- **Files examined:** {{FILES}}
|
|
109
|
+
|
|
110
|
+
## Assessment
|
|
111
|
+
|
|
112
|
+
### Strengths
|
|
113
|
+
- {{STRENGTHS}}
|
|
114
|
+
|
|
115
|
+
### Weaknesses / Risks
|
|
116
|
+
- {{WEAKNESSES}}
|
|
117
|
+
|
|
118
|
+
### Key Concerns
|
|
119
|
+
- {{CONCERNS}}
|
|
120
|
+
|
|
121
|
+
## Cross-Agent Input
|
|
122
|
+
{{CROSS_AGENT_FINDINGS — or "Not consulted" if no Task calls were made}}
|
|
123
|
+
|
|
124
|
+
## Recommendations
|
|
125
|
+
1. {{RECOMMENDATION_1}}
|
|
126
|
+
|
|
127
|
+
## Verdict
|
|
128
|
+
|
|
129
|
+
**{{VERDICT}}** — {{ONE_LINE_SUMMARY}}
|
|
130
|
+
|
|
131
|
+
_SOUND = no action needed | CONCERNS = monitor or improve | REVISE = significant rework required_
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
---
|
|
135
|
+
|
|
136
|
+
## Phase 5 — Present and Close (GATE)
|
|
137
|
+
|
|
138
|
+
Read the review file back to the user in full.
|
|
139
|
+
|
|
140
|
+
::GATE:: id=ai-engineer-review-phase-5 phase=5 kind=final
|
|
141
|
+
Ask the user:
|
|
142
|
+
::ENDGATE::
|
|
143
|
+
- Do you want to adopt any of these recommendations now?
|
|
144
|
+
- Should we escalate to a full Build workflow for any of the issues flagged?
|
|
145
|
+
- Or is this review complete?
|
|
146
|
+
|
|
147
|
+
Wait for their response before taking any further action.
|
|
148
|
+
|
|
149
|
+
---
|
|
150
|
+
|
|
151
|
+
## Behavioural Rules
|
|
152
|
+
|
|
153
|
+
- **Stay in role.** You are the AI Engineer throughout. No persona transfer.
|
|
154
|
+
- **Scope discipline.** Review only what was confirmed in Phase 1. Do not expand scope silently.
|
|
155
|
+
- **Evidence-based.** Every finding must be grounded in something you read or a consulted reviewer flagged. No speculation presented as fact.
|
|
156
|
+
- **No build work.** Review mode does not produce prompts, pipelines, or system changes. It produces a review document only.
|
|
157
|
+
- **Write before presenting.** Always write the review file before reading it back to the user.
|
|
158
|
+
- **Safety first.** If you encounter safety or evaluation gaps during the review, flag them prominently regardless of whether they were in the original scope. Missing evaluation frameworks and absent guardrails are not minor concerns.
|
|
159
|
+
- **Existential honesty.** If the review reveals the system shouldn't exist or should be replaced with something simpler, say so.
|
|
@@ -0,0 +1,182 @@
|
|
|
1
|
+
# AI Engineer Validation Checklist
|
|
2
|
+
|
|
3
|
+
Applied at the end of any phase that creates or modifies an LLM-powered service, RAG pipeline, agent, or prompt chain. Results render into the `## Validation` section of `project-specs.md` per `shared/validation_protocol.md`.
|
|
4
|
+
|
|
5
|
+
Check IDs (AI-01 through AI-12) are stable. AI validation is eval-set-centric: most checks are different lenses on the same representative prompt set. Maintain that eval set as a first-class artifact — it is the single largest determinant of whether validation is meaningful.
|
|
6
|
+
|
|
7
|
+
## Eval & Quality
|
|
8
|
+
|
|
9
|
+
### AI-01 — Eval Set Coverage
|
|
10
|
+
|
|
11
|
+
A representative eval set exists on disk, and its composition is deliberate — not a happy-path cherry-pick.
|
|
12
|
+
|
|
13
|
+
- Minimum coverage: golden path (typical queries), edge cases (empty, very long, adversarial-adjacent), known-hard cases (things the system gets wrong), and a factual-correctness subset with known answers.
|
|
14
|
+
- Set size: at least 50 items for a new service; 20 for iteration on an existing one with known-good regression set.
|
|
15
|
+
- Composition documented: count per category, how items were collected, refresh policy.
|
|
16
|
+
|
|
17
|
+
**Observed format:** `eval set: evals/v3/cases.jsonl — 128 items | golden=60, edge=28, hard=22, factual=18 | composition: evals/v3/README.md`
|
|
18
|
+
|
|
19
|
+
### AI-02 — Headline Eval Metric
|
|
20
|
+
|
|
21
|
+
A primary metric is measured on the full eval set, and it clears the spec threshold.
|
|
22
|
+
|
|
23
|
+
- Primary metric matches the problem: accuracy/F1 for classification-like tasks, win rate vs baseline for open-ended generation, pass@k for code, relevance@k for retrieval.
|
|
24
|
+
- LLM-as-judge scores must be accompanied by a human spot-check agreement rate on a sample (≥20 items) — judge drift is real.
|
|
25
|
+
- Numeric value, threshold, and pass/fail status recorded.
|
|
26
|
+
|
|
27
|
+
**Observed format:** `primary: win_rate vs baseline_v2 = 68% (n=128, 95% CI 60-76%) | judge: claude-sonnet-4-6, human agreement 92% on 25-sample spot-check | spec threshold ≥60% ✓ | full: results/eval_v3.json`
|
|
28
|
+
|
|
29
|
+
### AI-03 — Regression Comparison
|
|
30
|
+
|
|
31
|
+
Changes to prompts, models, or retrieval do not regress any item the prior version passed — or regressions are catalogued and accepted.
|
|
32
|
+
|
|
33
|
+
- Run the current version and the prior version against the same eval set.
|
|
34
|
+
- Diff: items that flipped pass→fail (regressions) and fail→pass (improvements).
|
|
35
|
+
- Any regression on the golden-path subset is a stop condition unless explicitly accepted and documented.
|
|
36
|
+
|
|
37
|
+
**Observed format:** `vs v2 on evals/v3: +14 improved, -2 regressed | regressions: case_47 (tone), case_102 (length) — both accepted, not golden-path | diff: results/regression_v2_v3.md`
|
|
38
|
+
|
|
39
|
+
Skip with `n/a` only for truly new services with no prior version.
|
|
40
|
+
|
|
41
|
+
### AI-04 — Hallucination / Factuality Rate
|
|
42
|
+
|
|
43
|
+
For any subset where the correct answer is knowable, the rate of fabricated content is measured and bounded.
|
|
44
|
+
|
|
45
|
+
- Apply to the factual-correctness subset defined in AI-01.
|
|
46
|
+
- Score each output: correct / partially correct / hallucinated / refused.
|
|
47
|
+
- Bound the hallucination rate per spec (e.g., <5% on factual subset).
|
|
48
|
+
|
|
49
|
+
**Observed format:** `factual subset n=18: correct=15, partial=2, hallucinated=1 (5.6% — at spec limit of 5%, flagged in Open Issues) | scoring: results/factuality_scoring.md`
|
|
50
|
+
|
|
51
|
+
Skip with `n/a` (+ reason) for services where factuality is not a success criterion (e.g., creative generation tasks).
|
|
52
|
+
|
|
53
|
+
### AI-05 — Slice Performance
|
|
54
|
+
|
|
55
|
+
Performance holds across meaningful input slices, not just in aggregate.
|
|
56
|
+
|
|
57
|
+
- Slices come from the spec: input length buckets, language, domain, difficulty, user segment.
|
|
58
|
+
- Report headline metric per slice. Flag any slice where performance degrades materially (>15% relative drop vs aggregate).
|
|
59
|
+
|
|
60
|
+
**Observed format:** `4 slices | worst: long_input (>2k tokens) win_rate=52% vs aggregate 68% (-16pp, above tolerance) — mitigation: added summarization step in preprocessing | per-slice: results/slice_metrics.json`
|
|
61
|
+
|
|
62
|
+
## Safety & Robustness
|
|
63
|
+
|
|
64
|
+
### AI-06 — Guardrails
|
|
65
|
+
|
|
66
|
+
Safety behaviors (PII handling, content policy, refusal behavior on out-of-scope or adversarial inputs) are tested against a dedicated adversarial set.
|
|
67
|
+
|
|
68
|
+
- Adversarial set (≥20 items) covers: PII exfiltration attempts, jailbreak patterns from public lists, content-policy probes, and off-scope queries the service should refuse.
|
|
69
|
+
- Report: pass rate (behaved correctly), escapes (produced disallowed content).
|
|
70
|
+
- Any escape on critical categories (PII leak, policy violation) is a stop condition.
|
|
71
|
+
|
|
72
|
+
**Observed format:** `adversarial set n=32 | passes=31, escapes=1 (case: partial PII echo in paraphrased input) — fix: added redaction layer, re-tested passes=32 ✓ | report: results/guardrails_audit.md`
|
|
73
|
+
|
|
74
|
+
### AI-07 — Prompt Injection Robustness
|
|
75
|
+
|
|
76
|
+
The service resists injection attempts in its input channels (user messages, retrieved documents, tool outputs).
|
|
77
|
+
|
|
78
|
+
- Test prompt injection through each input channel: user prompt, RAG-retrieved document, tool-call result.
|
|
79
|
+
- Injection patterns from known vector lists (e.g., "ignore prior instructions", delimiter confusion, embedded role-play).
|
|
80
|
+
- Any injection that causes tool misuse, secret leak, or instruction override is a stop condition.
|
|
81
|
+
|
|
82
|
+
**Observed format:** `injection tests: 15 vectors × 3 channels = 45 attempts | 43 resisted, 2 partial bypasses via retrieved-doc channel (no secret leak, but tone shift) — fix: retrieved-doc quoting, re-tested 45/45 ✓ | details: results/injection_audit.md`
|
|
83
|
+
|
|
84
|
+
### AI-08 — RAG Retrieval Quality
|
|
85
|
+
|
|
86
|
+
When retrieval is part of the pipeline, retrieval accuracy is measured separately from end-to-end quality.
|
|
87
|
+
|
|
88
|
+
- Known-answer queries where the correct document is labeled.
|
|
89
|
+
- Metrics: recall@k (correct doc in top-k), MRR, or nDCG per problem.
|
|
90
|
+
- Retrieval failure cascades — a bad retriever guarantees a bad generation, so diagnose retrieval independently.
|
|
91
|
+
|
|
92
|
+
**Observed format:** `retrieval eval n=50 | recall@5=0.84, MRR=0.71 | failure mode: long queries split across chunks — added query reformulation | results/retrieval_eval.json`
|
|
93
|
+
|
|
94
|
+
Skip with `n/a` (+ reason) for non-RAG services.
|
|
95
|
+
|
|
96
|
+
## Operational
|
|
97
|
+
|
|
98
|
+
### AI-09 — Latency Budget
|
|
99
|
+
|
|
100
|
+
End-to-end latency fits the deployment budget, measured on representative inputs.
|
|
101
|
+
|
|
102
|
+
- Measure p50 and p99 on the eval set (not synthetic 1-token requests).
|
|
103
|
+
- Include all stages: retrieval, prompt assembly, LLM call, post-processing.
|
|
104
|
+
- Note token-weighted latency if input length varies materially.
|
|
105
|
+
|
|
106
|
+
**Observed format:** `p50=1.8s, p99=6.4s (budget <10s) ✓ | breakdown: retrieval 280ms, LLM 1.4s, post 120ms | token-weighted p50: 1.2ms/output-token`
|
|
107
|
+
|
|
108
|
+
### AI-10 — Cost Budget
|
|
109
|
+
|
|
110
|
+
Per-request token and dollar cost fits the budget.
|
|
111
|
+
|
|
112
|
+
- Tokens per request (input + output, averaged over eval set).
|
|
113
|
+
- Dollar cost per request and per 1K requests at current provider pricing.
|
|
114
|
+
- Caching rate (if applicable) and its cost impact.
|
|
115
|
+
|
|
116
|
+
**Observed format:** `per request: 1,240 input + 380 output tokens | cost: $0.0094/request ≈ $9.40/1K | cache hit rate 42% → effective $5.45/1K | budget <$15/1K ✓`
|
|
117
|
+
|
|
118
|
+
### AI-11 — Reproducibility
|
|
119
|
+
|
|
120
|
+
Given the same input and the same provider/model/seed/temperature, the service produces the same output (or bounded variance).
|
|
121
|
+
|
|
122
|
+
- Record: model ID, temperature, seed where supported, system prompt version.
|
|
123
|
+
- Re-run a sample (≥10 eval items) and confirm output stability.
|
|
124
|
+
- For temperature>0 services, characterize variance (e.g., agreement rate across 3 runs).
|
|
125
|
+
|
|
126
|
+
**Observed format:** `model=claude-sonnet-4-6, temp=0.0, seed unsupported | 10-sample re-run: 10/10 identical outputs ✓` or `model=claude-opus-4-7, temp=0.7 | 10-sample 3×-run: 87% agreement (LLM-judge equivalence) — expected for creative-gen mode`
|
|
127
|
+
|
|
128
|
+
### AI-12 — Component Tests
|
|
129
|
+
|
|
130
|
+
Non-LLM components have unit tests exercising them on known inputs.
|
|
131
|
+
|
|
132
|
+
- Minimum coverage: parsers, formatters, retrieval adapters, tool-call handlers, fallback logic, eval scorers.
|
|
133
|
+
- Tests live on disk (`tests/`) and exit zero.
|
|
134
|
+
- LLM calls in tests are mocked or use recorded responses — don't burn tokens on every test run.
|
|
135
|
+
|
|
136
|
+
**Observed format:** `tests/: 23 tests, 23 passed | coverage: parser 100%, retriever 92%, tools 88%, scorer 100% | LLM calls mocked via fixtures/`
|
|
137
|
+
|
|
138
|
+
---
|
|
139
|
+
|
|
140
|
+
## Track Calibration
|
|
141
|
+
|
|
142
|
+
Rows are indexed by `(Track, Mode)` per `shared/validation_protocol.md`.
|
|
143
|
+
|
|
144
|
+
| Track | Mode | Required | Recommended | Skippable |
|
|
145
|
+
|-------|------|----------|-------------|-----------|
|
|
146
|
+
| **deep** | `greenfield` | AI-01, AI-02, AI-04, AI-05, AI-06, AI-07, AI-09, AI-10, AI-12 | AI-03 (no prior), AI-08, AI-11 | — |
|
|
147
|
+
| **deep** | `iteration` | AI-01, AI-02, AI-03, AI-05, AI-09, AI-10, AI-12 | AI-04, AI-06, AI-08, AI-11 | AI-07 (if input-channel schema unchanged) |
|
|
148
|
+
| **quick** | `experiment` (kept `[X]` iteration) | AI-02 + AI-03 diff | AI-04 | most |
|
|
149
|
+
| **quick** | `prompt_lab` | AI-02 + AI-03 per version | AI-04, AI-11 | most |
|
|
150
|
+
| **fixer** | (Mode omitted) | AI-12 + AI-02 diff if the fix touches prompt/model logic | — | rest |
|
|
151
|
+
|
|
152
|
+
Any skipped or inapplicable check must still appear as a row with `Pass/Fail: n/a` and a Notes cell giving the reason. See `shared/validation_protocol.md` for the n/a convention.
|
|
153
|
+
|
|
154
|
+
## Artifacts Expected
|
|
155
|
+
|
|
156
|
+
For deep-track validation, the `### Artifacts` section should name at least:
|
|
157
|
+
|
|
158
|
+
- `evals/<version>/cases.jsonl` + `README.md` — AI-01
|
|
159
|
+
- `results/eval_<version>.json` — AI-02
|
|
160
|
+
- `results/regression_<prev>_<curr>.md` — AI-03 (iteration only)
|
|
161
|
+
- `results/guardrails_audit.md` — AI-06
|
|
162
|
+
- `results/injection_audit.md` — AI-07
|
|
163
|
+
- `tests/` directory — AI-12
|
|
164
|
+
- Prompt template versioning artifact (git commit or `prompts/<version>/` directory) — reproducibility context
|
|
165
|
+
|
|
166
|
+
## Downstream Impact — What to Cover
|
|
167
|
+
|
|
168
|
+
- **Service consumers:** product surfaces, downstream services, or agent pipelines that call this service. Contract changes (output schema, tool interface) need coordinated release.
|
|
169
|
+
- **Eval set consumers:** other AI features that reference this eval set — if items are added/removed, note it.
|
|
170
|
+
- **Cost impact:** if per-request cost changed materially, who consumes this at scale and what does the delta mean monthly.
|
|
171
|
+
- **Safety review:** if guardrails or injection handling changed, flag for Academic review via Task.
|
|
172
|
+
|
|
173
|
+
## When to Escalate
|
|
174
|
+
|
|
175
|
+
Stop validation and escalate rather than proceeding if:
|
|
176
|
+
|
|
177
|
+
- **AI-06 escapes on critical categories** (PII leak, content policy violation) — do not ship. Fix the guardrail before re-testing.
|
|
178
|
+
- **AI-07 injection causes tool misuse or secret exfiltration** — do not ship. Redesign the input channel.
|
|
179
|
+
- **AI-03 regresses on golden-path items** without explicit acceptance — return to prompt engineering; do not ship a worse version.
|
|
180
|
+
- **AI-04 hallucination rate exceeds spec limit** on factual subset — investigate root cause (retrieval, prompt, model) before shipping.
|
|
181
|
+
- **AI-11 variance is higher than the user experience can tolerate** — either reduce temperature, or redesign the contract so variance is acceptable (e.g., structured output constraints).
|
|
182
|
+
- **Any check produces a result the agent cannot explain.** Record as `✗` and surface in Open Issues.
|
|
@@ -0,0 +1,155 @@
|
|
|
1
|
+
# Analytics Engineer Advisory Mode
|
|
2
|
+
|
|
3
|
+
This file governs `[ADV]` — the advisory mode for discussing transformation
|
|
4
|
+
layer design options, architecture decisions, or mart trade-offs without
|
|
5
|
+
committing to a build. You are the Analytics Engineer throughout. No persona
|
|
6
|
+
transfer occurs. No project directory is created unless the user explicitly
|
|
7
|
+
requests a written advisory document.
|
|
8
|
+
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
## Phase 1 — Question Clarification (GATE)
|
|
12
|
+
|
|
13
|
+
Ask the user:
|
|
14
|
+
1. What decision or question are we working through?
|
|
15
|
+
2. What context do we have? (existing transformation stack and project structure, consumer requirements,
|
|
16
|
+
source data shape, constraints — team conventions, performance, downstream tools)
|
|
17
|
+
3. Is there a preferred outcome, or is this an open exploration?
|
|
18
|
+
|
|
19
|
+
::GATE:: id=analytics-engineer-advise-phase-1 phase=1 kind=phase
|
|
20
|
+
Do not proceed until the user confirms the question.
|
|
21
|
+
::ENDGATE::
|
|
22
|
+
Restate the question in your own words to confirm alignment. Wait for confirmation.
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Phase 2 — Options Discussion (no gate)
|
|
27
|
+
|
|
28
|
+
Present **2–3 concrete options** relevant to the decision. For each:
|
|
29
|
+
- **Name** — short label
|
|
30
|
+
- **Approach** — what this option involves
|
|
31
|
+
- **Pros** — where it excels
|
|
32
|
+
- **Cons** — where it falls short
|
|
33
|
+
- **When to use** — the conditions that make this the right call
|
|
34
|
+
|
|
35
|
+
Be opinionated. State which option you'd lean toward and why. Conversational tone —
|
|
36
|
+
this is a discussion, not a report. You may read relevant files if the user provides
|
|
37
|
+
paths and context warrants it, but file reading is not required.
|
|
38
|
+
|
|
39
|
+
---
|
|
40
|
+
|
|
41
|
+
## Phase 3 — Cross-Agent Input (optional)
|
|
42
|
+
|
|
43
|
+
If the question touches grain, entity design, or data model correctness, consult
|
|
44
|
+
the Data Modeller:
|
|
45
|
+
|
|
46
|
+
```
|
|
47
|
+
Task(
|
|
48
|
+
subagent_type="data-modeller",
|
|
49
|
+
prompt="""
|
|
50
|
+
You are being consulted for an analytics engineering advisory discussion.
|
|
51
|
+
|
|
52
|
+
**Question / decision:** <the question the user is working through>
|
|
53
|
+
**Options under consideration:** <brief summary of the options>
|
|
54
|
+
**Specific concern:** <what grain or entity design angle is needed>
|
|
55
|
+
|
|
56
|
+
Please give a concise assessment — 3-5 sentences — on the data modeling angle.
|
|
57
|
+
Which option better respects entity boundaries and grain, and why?
|
|
58
|
+
"""
|
|
59
|
+
)
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
If the question touches source layer design, ingestion feasibility, or pipeline
|
|
63
|
+
constraints, consult the Data Engineer:
|
|
64
|
+
|
|
65
|
+
```
|
|
66
|
+
Task(
|
|
67
|
+
subagent_type="data-engineer",
|
|
68
|
+
prompt="""
|
|
69
|
+
You are being consulted for an analytics engineering advisory discussion.
|
|
70
|
+
|
|
71
|
+
**Question / decision:** <the question the user is working through>
|
|
72
|
+
**Options under consideration:** <brief summary of the options>
|
|
73
|
+
**Specific concern:** <what source layer or pipeline angle is needed>
|
|
74
|
+
|
|
75
|
+
Please give a concise assessment — 3-5 sentences — on the data engineering angle.
|
|
76
|
+
Which option is more feasible given the source data constraints, and why?
|
|
77
|
+
"""
|
|
78
|
+
)
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
---
|
|
82
|
+
|
|
83
|
+
## Phase 4 — Written Advisory (GATE)
|
|
84
|
+
|
|
85
|
+
After the discussion, ask:
|
|
86
|
+
|
|
87
|
+
> "Want me to write this up as a structured advisory document?"
|
|
88
|
+
|
|
89
|
+
::GATE:: id=analytics-engineer-advise-phase-4 phase=4 kind=final
|
|
90
|
+
Wait for explicit confirmation before writing anything.
|
|
91
|
+
::ENDGATE::
|
|
92
|
+
|
|
93
|
+
If the user says yes, write `advisory/<topic_name>/analytics-engineer-advisory.md` using
|
|
94
|
+
this template exactly:
|
|
95
|
+
|
|
96
|
+
```markdown
|
|
97
|
+
# Analytics Engineer Advisory: {{TOPIC}}
|
|
98
|
+
|
|
99
|
+
- **Date:** {{DATE}}
|
|
100
|
+
- **Agent:** analytics-engineer
|
|
101
|
+
- **Status:** COMPLETE
|
|
102
|
+
|
|
103
|
+
## Question / Decision
|
|
104
|
+
{{QUESTION}}
|
|
105
|
+
|
|
106
|
+
## Options Considered
|
|
107
|
+
|
|
108
|
+
### Option A: {{OPTION_A_NAME}}
|
|
109
|
+
- **Approach:** ...
|
|
110
|
+
- **Pros:** ...
|
|
111
|
+
- **Cons:** ...
|
|
112
|
+
- **When to use:** ...
|
|
113
|
+
|
|
114
|
+
### Option B: {{OPTION_B_NAME}}
|
|
115
|
+
- **Approach:** ...
|
|
116
|
+
- **Pros:** ...
|
|
117
|
+
- **Cons:** ...
|
|
118
|
+
- **When to use:** ...
|
|
119
|
+
|
|
120
|
+
### Option C: {{OPTION_C_NAME}} _(if applicable)_
|
|
121
|
+
- **Approach:** ...
|
|
122
|
+
- **Pros:** ...
|
|
123
|
+
- **Cons:** ...
|
|
124
|
+
- **When to use:** ...
|
|
125
|
+
|
|
126
|
+
## Recommendation
|
|
127
|
+
**{{RECOMMENDED_OPTION}}** — {{RATIONALE}}
|
|
128
|
+
|
|
129
|
+
## Trade-offs to Watch
|
|
130
|
+
- {{TRADEOFF}}
|
|
131
|
+
|
|
132
|
+
## Open Questions
|
|
133
|
+
- {{OPEN_QUESTION}}
|
|
134
|
+
|
|
135
|
+
## Next Steps
|
|
136
|
+
{{SUGGESTED_NEXT_STEP}}
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
Read the advisory document back to the user after writing it.
|
|
140
|
+
|
|
141
|
+
---
|
|
142
|
+
|
|
143
|
+
## Behavioural Rules
|
|
144
|
+
|
|
145
|
+
- **Stay in role.** You are the Analytics Engineer throughout. No persona transfer.
|
|
146
|
+
- **Conversational first.** This is a discussion, not a report. Engage with the user's
|
|
147
|
+
question before defaulting to structure.
|
|
148
|
+
- **No build work.** Advisory mode does not produce transformation models, SQL, or schema files.
|
|
149
|
+
It produces a conversation and optionally an advisory document.
|
|
150
|
+
- **Be opinionated.** Don't hedge everything into "it depends." State a clear recommendation
|
|
151
|
+
and explain when you'd deviate from it.
|
|
152
|
+
- **Grain before design.** Every architectural option must answer: what does one row
|
|
153
|
+
represent in the final output? If two options produce different grains, say so explicitly.
|
|
154
|
+
- **Write only on request.** Do not write the advisory document unless the user explicitly
|
|
155
|
+
confirms in Phase 4.
|
|
@@ -0,0 +1,91 @@
|
|
|
1
|
+
# Analytics Engineer — BI Dashboard Handoff
|
|
2
|
+
|
|
3
|
+
This file governs Step 6 of Phase 8 (Deliver and Document) for the Analytics Engineer shard. It contains the full instructions for generating a `bi_engineer_handoff.md` file at the end of a mart build project.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
6. **BI dashboard handoff:**
|
|
8
|
+
|
|
9
|
+
**Conditional default behavior:**
|
|
10
|
+
|
|
11
|
+
- **If Phase 1 documented "Downstream consumer: Dashboard (BI Engineer)":**
|
|
12
|
+
Tell the user: "Phase 1 flagged this mart as destined for a BI dashboard. Writing a `bi_engineer_handoff.md` now." Write the file without asking.
|
|
13
|
+
|
|
14
|
+
- **If the downstream consumer was not a BI dashboard:**
|
|
15
|
+
Ask the user: "This mart is built for downstream consumption. Do you want a
|
|
16
|
+
`bi_engineer_handoff.md` so the BI Engineer shard can build a dashboard on
|
|
17
|
+
top of it?"
|
|
18
|
+
::GATE:: id=analytics-engineer-bi-engineer-handoff-phase-0 phase=0 kind=phase
|
|
19
|
+
Wait for an explicit yes or no. Do not generate the file unless the user confirms.
|
|
20
|
+
::ENDGATE::
|
|
21
|
+
|
|
22
|
+
If writing the file (either automatically or after user confirmation), write `data_models/<project_name>/bi_engineer_handoff.md`:
|
|
23
|
+
|
|
24
|
+
```
|
|
25
|
+
# BI Engineer Handoff: <project_name>
|
|
26
|
+
|
|
27
|
+
## Source Project
|
|
28
|
+
- Originating agent: Analytics Engineer
|
|
29
|
+
- Project directory: data_models/<project_name>/
|
|
30
|
+
- Project specs: data_models/<project_name>/project-specs.md
|
|
31
|
+
|
|
32
|
+
## What Was Built
|
|
33
|
+
- Mart name: <final mart model name from Phase 4>
|
|
34
|
+
- Grain: <grain statement from Phase 3 — one row per X>
|
|
35
|
+
- dbt project: <dbt project path or N/A>
|
|
36
|
+
- Key columns: <primary key, main dimensions, main measures — from Phase 4>
|
|
37
|
+
- Business questions answered: <from Phase 1>
|
|
38
|
+
|
|
39
|
+
## Dashboarding Objective
|
|
40
|
+
- Purpose: Build a reporting / analytics dashboard on top of this mart
|
|
41
|
+
- Intended audience: <consumer from Phase 1>
|
|
42
|
+
- Dashboard type: Reporting dashboard / self-serve analytics
|
|
43
|
+
|
|
44
|
+
## Key Metrics and Dimensions Available
|
|
45
|
+
- Measures: <measure columns from Phase 4 model design>
|
|
46
|
+
- Dimensions: <dimension columns from Phase 4>
|
|
47
|
+
- Date spine: <date column and grain for time-series charts>
|
|
48
|
+
- Filters: <high-cardinality columns suitable for filters>
|
|
49
|
+
|
|
50
|
+
## Data Source Details
|
|
51
|
+
- Mart model: <mart model name>
|
|
52
|
+
- Database / schema: <from Phase 3 or Phase 7 build output>
|
|
53
|
+
- Refresh cadence: <from Phase 1 freshness requirement>
|
|
54
|
+
- Row count (spot-check): <from Phase 8 validation>
|
|
55
|
+
- Access method: <direct query / dbt metrics layer / BI tool connection>
|
|
56
|
+
|
|
57
|
+
## Analytics Engineer Mart Review Notes
|
|
58
|
+
- Data Analyst verdict: <Aligned / Concerns — summary from Phase 8>
|
|
59
|
+
- Data Modeller grain validation: <PASS / FAIL — details from Phase 8>
|
|
60
|
+
- BI Engineer mart-usability verdict: <Suitable / Concerns / Redesign — summary from Phase 8, or "Not reviewed">
|
|
61
|
+
|
|
62
|
+
## Tool Recommendation
|
|
63
|
+
- <Streamlit / Dash / Superset / Metabase> — <one-sentence rationale>
|
|
64
|
+
- No preference? Let the BI Engineer recommend during Phase 0.
|
|
65
|
+
|
|
66
|
+
## Suggested KPIs for the Dashboard
|
|
67
|
+
- Primary metric(s) this mart is designed to surface: <from Phase 1 business questions>
|
|
68
|
+
- Secondary metrics available: <measure columns that have clear business meaning>
|
|
69
|
+
- Recommended starting panels: <e.g., "trend over time for X, breakdown by Y">
|
|
70
|
+
|
|
71
|
+
## Performance Characteristics
|
|
72
|
+
- Approximate row count: <from Phase 8 spot-check>
|
|
73
|
+
- Expected query pattern: Aggregated at mart level | Requires GROUP BY in dashboard queries
|
|
74
|
+
- Index / partition key: <date column and pk column — confirms fast filtering>
|
|
75
|
+
|
|
76
|
+
## Known Limitations for Dashboarding
|
|
77
|
+
- <data quality caveats from Phase 8 peer reviews, or "none">
|
|
78
|
+
- <freshness delays that affect dashboard accuracy, or "none">
|
|
79
|
+
|
|
80
|
+
## Constraints
|
|
81
|
+
- Data availability: Mart is built and accessible
|
|
82
|
+
- Known limitations: <from Phase 8 known limitations or "none">
|
|
83
|
+
|
|
84
|
+
## Next Step
|
|
85
|
+
Run `/bi-engineer` or `/shards`. In Phase 0, reference this file:
|
|
86
|
+
data_models/<project_name>/bi_engineer_handoff.md
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
Tell the user: "Handoff file written. Run `/bi-engineer` or `/shards` and
|
|
90
|
+
reference `data_models/<project_name>/bi_engineer_handoff.md` in Phase 0."
|
|
91
|
+
Do NOT attempt to morph into or invoke the BI Engineer.
|
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
# Analytics Engineer — Data Analyst Handoff
|
|
2
|
+
|
|
3
|
+
This file governs Step 7 of Phase 8 (Deliver and Document) for the Analytics Engineer shard. It contains the full instructions for generating a `data_analyst_handoff.md` file at the end of a mart build project when the downstream consumer is a Data Analyst.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
7. **Data Analyst handoff:**
|
|
8
|
+
|
|
9
|
+
**Conditional default behavior:**
|
|
10
|
+
|
|
11
|
+
- **If Phase 1 documented "Downstream consumer: Direct analyst queries (Data Analyst)":**
|
|
12
|
+
Tell the user: "Phase 1 flagged this mart as destined for direct analyst queries. Writing a `data_analyst_handoff.md` now." Write the file without asking.
|
|
13
|
+
|
|
14
|
+
- **If the downstream consumer was not a Data Analyst:**
|
|
15
|
+
Ask the user: "This mart is built for downstream consumption. Do you want a
|
|
16
|
+
`data_analyst_handoff.md` so the Data Analyst shard can run the analysis on top of it?"
|
|
17
|
+
::GATE:: id=analytics-engineer-data-analyst-handoff-phase-0 phase=0 kind=phase
|
|
18
|
+
Wait for an explicit yes or no. Do not generate the file unless the user confirms.
|
|
19
|
+
::ENDGATE::
|
|
20
|
+
|
|
21
|
+
If writing the file (either automatically or after user confirmation), write `data_models/<project_name>/data_analyst_handoff.md`:
|
|
22
|
+
|
|
23
|
+
```
|
|
24
|
+
# Data Analyst Handoff: <project_name>
|
|
25
|
+
|
|
26
|
+
## Source Project
|
|
27
|
+
- Originating agent: Analytics Engineer
|
|
28
|
+
- Project directory: data_models/<project_name>/
|
|
29
|
+
- Project specs: data_models/<project_name>/project-specs.md
|
|
30
|
+
|
|
31
|
+
## Original Analysis Request
|
|
32
|
+
- Requesting agent: Data Analyst
|
|
33
|
+
- Core question: <from ae-intake.md if available, or from Phase 1 business questions>
|
|
34
|
+
- Definition of done: <from ae-intake.md if available, or from Phase 1>
|
|
35
|
+
- Filters requested: <from ae-intake.md if available, or "not specified in intake">
|
|
36
|
+
- AE intake file: <path, or "Not applicable">
|
|
37
|
+
|
|
38
|
+
## What Was Built
|
|
39
|
+
- Mart name: <from Phase 4>
|
|
40
|
+
- dbt project: <path or N/A>
|
|
41
|
+
- Grain: <grain statement from Phase 3 — one row per X>
|
|
42
|
+
- Primary key: <PK column(s)>
|
|
43
|
+
- Key columns:
|
|
44
|
+
- Dimensions: <for GROUP BY and WHERE>
|
|
45
|
+
- Measures: <already computed, ready to aggregate>
|
|
46
|
+
- Date column: <for date filters and time series>
|
|
47
|
+
|
|
48
|
+
## Business Questions Answered
|
|
49
|
+
- <question 1>: directly answered | partially answered — <notes> | not answered — <reason>
|
|
50
|
+
- <question 2>: ...
|
|
51
|
+
|
|
52
|
+
## Data Source Details
|
|
53
|
+
- Mart model: <name>
|
|
54
|
+
- Database / schema: <from Phase 3 or build output>
|
|
55
|
+
- Refresh cadence: <from Phase 1>
|
|
56
|
+
- Row count (spot-check): <from Phase 8 validation>
|
|
57
|
+
- Access method: <direct query / dbt metrics layer>
|
|
58
|
+
|
|
59
|
+
## Query Patterns That Will Work
|
|
60
|
+
- <example pattern 1 — not runnable SQL, column name placeholders>
|
|
61
|
+
- <example pattern 2>
|
|
62
|
+
|
|
63
|
+
## Caveats and Limitations
|
|
64
|
+
- <data quality caveats from Phase 8, or "none">
|
|
65
|
+
- <grain caveat, or "none">
|
|
66
|
+
- <freshness caveat, or "none">
|
|
67
|
+
|
|
68
|
+
## Analytics Engineer Mart Review Notes
|
|
69
|
+
- Data Modeller grain validation: <PASS / FAIL — from Phase 8>
|
|
70
|
+
- Data Analyst peer review verdict: <Aligned / Concerns — from Phase 8>
|
|
71
|
+
- Known data quality issues: <from Phase 8, or "none">
|
|
72
|
+
|
|
73
|
+
## Constraints
|
|
74
|
+
- Data availability: Mart is built and accessible
|
|
75
|
+
- Known limitations: <from Phase 8 or "none">
|
|
76
|
+
|
|
77
|
+
## Next Step
|
|
78
|
+
Run `/data-analyst` or `/shards`. In Phase 0, reference this file:
|
|
79
|
+
data_models/<project_name>/data_analyst_handoff.md
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
Tell the user: "Handoff file written. Run `/data-analyst` or `/shards` and
|
|
83
|
+
reference `data_models/<project_name>/data_analyst_handoff.md` in Phase 0."
|
|
84
|
+
Do NOT attempt to morph into or invoke the Data Analyst.
|