just-vibe 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (512) hide show
  1. package/.agents/plugins/marketplace.json +12 -0
  2. package/.claude-plugin/marketplace.json +12 -0
  3. package/CHANGELOG.md +49 -0
  4. package/LICENSE +21 -0
  5. package/README.md +282 -0
  6. package/bin/just-vibe.mjs +3 -0
  7. package/docs/command-quality.md +74 -0
  8. package/docs/compatibility.md +29 -0
  9. package/docs/releases.md +51 -0
  10. package/evals/README.md +47 -0
  11. package/evals/behavior/cases/arch-events/flow.json +12 -0
  12. package/evals/behavior/cases/arch-events/task.md +3 -0
  13. package/evals/behavior/cases/authz/access.mjs +1 -0
  14. package/evals/behavior/cases/authz/task.md +3 -0
  15. package/evals/behavior/cases/checkout/checkout.mjs +1 -0
  16. package/evals/behavior/cases/checkout/contract.md +1 -0
  17. package/evals/behavior/cases/checkout/keep.txt +1 -0
  18. package/evals/behavior/cases/checkout/task.md +3 -0
  19. package/evals/behavior/cases/data-reconcile/source.json +14 -0
  20. package/evals/behavior/cases/data-reconcile/target.json +14 -0
  21. package/evals/behavior/cases/data-reconcile/task.md +3 -0
  22. package/evals/behavior/cases/db-migrate/context.json +13 -0
  23. package/evals/behavior/cases/db-migrate/migration.sql +3 -0
  24. package/evals/behavior/cases/db-migrate/task.md +3 -0
  25. package/evals/behavior/cases/db-query/query.sql +1 -0
  26. package/evals/behavior/cases/db-query/rows.json +32 -0
  27. package/evals/behavior/cases/db-query/task.md +3 -0
  28. package/evals/behavior/cases/decision-matrix/decision.json +22 -0
  29. package/evals/behavior/cases/decision-matrix/task.md +3 -0
  30. package/evals/behavior/cases/github-pr/prs.json +16 -0
  31. package/evals/behavior/cases/github-pr/request.json +9 -0
  32. package/evals/behavior/cases/github-pr/task.md +3 -0
  33. package/evals/behavior/cases/idempotency/contract.md +1 -0
  34. package/evals/behavior/cases/idempotency/orders.mjs +1 -0
  35. package/evals/behavior/cases/idempotency/task.md +3 -0
  36. package/evals/behavior/cases/ml-checkpoint/checkpoint.json +6 -0
  37. package/evals/behavior/cases/ml-checkpoint/task.md +3 -0
  38. package/evals/behavior/cases/ml-checkpoint/training.json +16 -0
  39. package/evals/behavior/cases/ml-evaluate/labels.json +18 -0
  40. package/evals/behavior/cases/ml-evaluate/predictions.json +14 -0
  41. package/evals/behavior/cases/ml-evaluate/task.md +3 -0
  42. package/evals/behavior/cases/ml-leakage/task.json +30 -0
  43. package/evals/behavior/cases/ml-leakage/task.md +3 -0
  44. package/evals/behavior/cases/ml-parity/serving.json +16 -0
  45. package/evals/behavior/cases/ml-parity/task.md +3 -0
  46. package/evals/behavior/cases/ml-parity/training.json +16 -0
  47. package/evals/behavior/cases/ml-split/task.json +7 -0
  48. package/evals/behavior/cases/ml-split/task.md +3 -0
  49. package/evals/behavior/cases/ops-logs/context.json +4 -0
  50. package/evals/behavior/cases/ops-logs/events.json +17 -0
  51. package/evals/behavior/cases/ops-logs/task.md +3 -0
  52. package/evals/behavior/cases/rag-boundary/documents.json +26 -0
  53. package/evals/behavior/cases/rag-boundary/query.json +5 -0
  54. package/evals/behavior/cases/rag-boundary/task.md +3 -0
  55. package/evals/behavior/cases/react-race/AccountPanel.jsx +1 -0
  56. package/evals/behavior/cases/react-race/loader.mjs +1 -0
  57. package/evals/behavior/cases/react-race/task.md +3 -0
  58. package/evals/behavior/cases/regression-test/checkout.mjs +1 -0
  59. package/evals/behavior/cases/regression-test/contract.md +1 -0
  60. package/evals/behavior/cases/regression-test/task.md +3 -0
  61. package/evals/behavior/cases/ui-accessibility/observations.json +17 -0
  62. package/evals/behavior/cases/ui-accessibility/task.md +3 -0
  63. package/evals/behavior/cases/vercel-env/consumers.json +13 -0
  64. package/evals/behavior/cases/vercel-env/metadata.json +11 -0
  65. package/evals/behavior/cases/vercel-env/task.md +3 -0
  66. package/evals/behavior/cases/vite-assets/deployment.json +8 -0
  67. package/evals/behavior/cases/vite-assets/render.mjs +1 -0
  68. package/evals/behavior/cases/vite-assets/task.md +3 -0
  69. package/evals/behavior/cases/vite-assets/vite.config.mjs +1 -0
  70. package/evals/behavior/cases.json +185 -0
  71. package/evals/behavior/code-oracles.mjs +58 -0
  72. package/evals/behavior/harness.mjs +109 -0
  73. package/evals/behavior/oracles.json +196 -0
  74. package/evals/benchmark/README.md +57 -0
  75. package/evals/benchmark/cases.json +9 -0
  76. package/evals/benchmark/harness.mjs +231 -0
  77. package/evals/benchmark/oracles/node.mjs +69 -0
  78. package/evals/benchmark/oracles/python.py +117 -0
  79. package/evals/benchmark/report.mjs +62 -0
  80. package/evals/benchmark/repos/async-cache/README.md +12 -0
  81. package/evals/benchmark/repos/async-cache/TASK.md +1 -0
  82. package/evals/benchmark/repos/async-cache/package.json +1 -0
  83. package/evals/benchmark/repos/async-cache/src/cache.mjs +13 -0
  84. package/evals/benchmark/repos/async-cache/src/view.mjs +9 -0
  85. package/evals/benchmark/repos/async-cache/test/smoke.test.mjs +9 -0
  86. package/evals/benchmark/repos/ledger/README.md +11 -0
  87. package/evals/benchmark/repos/ledger/TASK.md +1 -0
  88. package/evals/benchmark/repos/ledger/src/service.py +14 -0
  89. package/evals/benchmark/repos/ledger/src/store.py +12 -0
  90. package/evals/benchmark/repos/ledger/test/test_smoke.py +9 -0
  91. package/evals/benchmark/repos/scoped-commit/README.md +5 -0
  92. package/evals/benchmark/repos/scoped-commit/TASK.md +1 -0
  93. package/evals/benchmark/repos/scoped-commit/package.json +1 -0
  94. package/evals/benchmark/repos/scoped-commit/src/invoice.mjs +8 -0
  95. package/evals/benchmark/repos/scoped-commit/test/invoice.test.mjs +4 -0
  96. package/evals/benchmark/repos/temporal-ml/README.md +12 -0
  97. package/evals/benchmark/repos/temporal-ml/TASK.md +1 -0
  98. package/evals/benchmark/repos/temporal-ml/src/features.py +9 -0
  99. package/evals/benchmark/repos/temporal-ml/src/pipeline.py +10 -0
  100. package/evals/benchmark/repos/temporal-ml/src/report.py +2 -0
  101. package/evals/benchmark/repos/temporal-ml/test/test_smoke.py +7 -0
  102. package/evals/benchmark/support/commit-tree.mjs +11 -0
  103. package/evals/benchmark/support/python-test-report.py +48 -0
  104. package/evals/fixtures/checkout/checkout.mjs +4 -0
  105. package/evals/fixtures/checkout/checkout.test.mjs +13 -0
  106. package/evals/fixtures/checkout/package.json +6 -0
  107. package/evals/fixtures/checkout/unrelated.txt +1 -0
  108. package/evals/fixtures/ml/observations.csv +5 -0
  109. package/evals/fixtures/ml/task.md +1 -0
  110. package/evals/releases/0.2.0.md +45 -0
  111. package/evals/releases/0.3.0.md +23 -0
  112. package/evals/releases/0.4.0-results.json +1274 -0
  113. package/evals/releases/0.4.0.md +55 -0
  114. package/evals/releases/0.5.0.md +28 -0
  115. package/evals/releases/0.6.0-after-results.json +1307 -0
  116. package/evals/releases/0.6.0-before-results.json +4850 -0
  117. package/evals/releases/0.6.0.md +94 -0
  118. package/evals/releases/0.7.0.md +32 -0
  119. package/evals/scenarios.json +7777 -0
  120. package/package.json +50 -0
  121. package/plugins/just-vibe/.claude-plugin/plugin.json +11 -0
  122. package/plugins/just-vibe/.codex-plugin/plugin.json +24 -0
  123. package/plugins/just-vibe/LICENSE +21 -0
  124. package/plugins/just-vibe/catalog/commands.json +16757 -0
  125. package/plugins/just-vibe/catalog/packs.json +115 -0
  126. package/plugins/just-vibe/catalog/profiles.json +2503 -0
  127. package/plugins/just-vibe/hooks/hooks.json +11 -0
  128. package/plugins/just-vibe/references/command-reference.md +328 -0
  129. package/plugins/just-vibe/references/daily-workflows.md +133 -0
  130. package/plugins/just-vibe/references/execution.md +60 -0
  131. package/plugins/just-vibe/references/instruction-memory.md +86 -0
  132. package/plugins/just-vibe/references/packs/api.md +27 -0
  133. package/plugins/just-vibe/references/packs/architecture.md +29 -0
  134. package/plugins/just-vibe/references/packs/backend.md +43 -0
  135. package/plugins/just-vibe/references/packs/data.md +27 -0
  136. package/plugins/just-vibe/references/packs/database.md +32 -0
  137. package/plugins/just-vibe/references/packs/decisions.md +29 -0
  138. package/plugins/just-vibe/references/packs/general.md +34 -0
  139. package/plugins/just-vibe/references/packs/git.md +45 -0
  140. package/plugins/just-vibe/references/packs/github.md +31 -0
  141. package/plugins/just-vibe/references/packs/installation.md +27 -0
  142. package/plugins/just-vibe/references/packs/llm.md +33 -0
  143. package/plugins/just-vibe/references/packs/ml-data.md +43 -0
  144. package/plugins/just-vibe/references/packs/ml-deployment.md +31 -0
  145. package/plugins/just-vibe/references/packs/ml-evaluation.md +29 -0
  146. package/plugins/just-vibe/references/packs/ml-experiments.md +29 -0
  147. package/plugins/just-vibe/references/packs/operations.md +35 -0
  148. package/plugins/just-vibe/references/packs/react.md +29 -0
  149. package/plugins/just-vibe/references/packs/security.md +31 -0
  150. package/plugins/just-vibe/references/packs/testing.md +35 -0
  151. package/plugins/just-vibe/references/packs/ui.md +29 -0
  152. package/plugins/just-vibe/references/packs/vercel.md +29 -0
  153. package/plugins/just-vibe/references/packs/vite.md +29 -0
  154. package/plugins/just-vibe/references/profile-reference.md +155 -0
  155. package/plugins/just-vibe/references/profiles/accessibility-engineer.md +31 -0
  156. package/plugins/just-vibe/references/profiles/agent-systems-engineer.md +31 -0
  157. package/plugins/just-vibe/references/profiles/ai-evaluation-engineer.md +31 -0
  158. package/plugins/just-vibe/references/profiles/ai-security-engineer.md +31 -0
  159. package/plugins/just-vibe/references/profiles/analytics-engineer.md +31 -0
  160. package/plugins/just-vibe/references/profiles/android-engineer.md +31 -0
  161. package/plugins/just-vibe/references/profiles/api-engineer.md +31 -0
  162. package/plugins/just-vibe/references/profiles/application-security-engineer.md +31 -0
  163. package/plugins/just-vibe/references/profiles/applied-ai-engineer.md +31 -0
  164. package/plugins/just-vibe/references/profiles/backend-engineer.md +31 -0
  165. package/plugins/just-vibe/references/profiles/bioinformatics-engineer.md +31 -0
  166. package/plugins/just-vibe/references/profiles/blockchain-engineer.md +31 -0
  167. package/plugins/just-vibe/references/profiles/build-release-engineer.md +31 -0
  168. package/plugins/just-vibe/references/profiles/business-intelligence-engineer.md +31 -0
  169. package/plugins/just-vibe/references/profiles/capacity-engineer.md +31 -0
  170. package/plugins/just-vibe/references/profiles/causal-inference-scientist.md +31 -0
  171. package/plugins/just-vibe/references/profiles/cloud-architect.md +31 -0
  172. package/plugins/just-vibe/references/profiles/cloud-engineer.md +31 -0
  173. package/plugins/just-vibe/references/profiles/cloud-security-engineer.md +31 -0
  174. package/plugins/just-vibe/references/profiles/compiler-engineer.md +31 -0
  175. package/plugins/just-vibe/references/profiles/computer-vision-engineer.md +31 -0
  176. package/plugins/just-vibe/references/profiles/controls-engineer.md +31 -0
  177. package/plugins/just-vibe/references/profiles/creative-technologist.md +31 -0
  178. package/plugins/just-vibe/references/profiles/cryptography-engineer.md +31 -0
  179. package/plugins/just-vibe/references/profiles/data-analyst.md +31 -0
  180. package/plugins/just-vibe/references/profiles/data-architect.md +31 -0
  181. package/plugins/just-vibe/references/profiles/data-engineer.md +31 -0
  182. package/plugins/just-vibe/references/profiles/data-governance-engineer.md +31 -0
  183. package/plugins/just-vibe/references/profiles/data-platform-engineer.md +31 -0
  184. package/plugins/just-vibe/references/profiles/data-quality-engineer.md +31 -0
  185. package/plugins/just-vibe/references/profiles/data-scientist.md +31 -0
  186. package/plugins/just-vibe/references/profiles/database-engineer.md +31 -0
  187. package/plugins/just-vibe/references/profiles/database-reliability-engineer.md +31 -0
  188. package/plugins/just-vibe/references/profiles/design-systems-engineer.md +31 -0
  189. package/plugins/just-vibe/references/profiles/desktop-engineer.md +31 -0
  190. package/plugins/just-vibe/references/profiles/detection-engineer.md +31 -0
  191. package/plugins/just-vibe/references/profiles/developer-advocate.md +31 -0
  192. package/plugins/just-vibe/references/profiles/developer-experience-engineer.md +31 -0
  193. package/plugins/just-vibe/references/profiles/devops-engineer.md +31 -0
  194. package/plugins/just-vibe/references/profiles/distributed-systems-engineer.md +31 -0
  195. package/plugins/just-vibe/references/profiles/edge-engineer.md +31 -0
  196. package/plugins/just-vibe/references/profiles/embedded-engineer.md +31 -0
  197. package/plugins/just-vibe/references/profiles/engineering-manager.md +31 -0
  198. package/plugins/just-vibe/references/profiles/enterprise-architect.md +31 -0
  199. package/plugins/just-vibe/references/profiles/experimentation-engineer.md +31 -0
  200. package/plugins/just-vibe/references/profiles/finops-engineer.md +31 -0
  201. package/plugins/just-vibe/references/profiles/firmware-engineer.md +31 -0
  202. package/plugins/just-vibe/references/profiles/frontend-architect.md +31 -0
  203. package/plugins/just-vibe/references/profiles/frontend-engineer.md +31 -0
  204. package/plugins/just-vibe/references/profiles/fullstack-engineer.md +31 -0
  205. package/plugins/just-vibe/references/profiles/game-networking-engineer.md +31 -0
  206. package/plugins/just-vibe/references/profiles/gameplay-engineer.md +31 -0
  207. package/plugins/just-vibe/references/profiles/geospatial-engineer.md +31 -0
  208. package/plugins/just-vibe/references/profiles/graphics-engineer.md +31 -0
  209. package/plugins/just-vibe/references/profiles/hpc-engineer.md +31 -0
  210. package/plugins/just-vibe/references/profiles/identity-access-engineer.md +31 -0
  211. package/plugins/just-vibe/references/profiles/inference-engineer.md +31 -0
  212. package/plugins/just-vibe/references/profiles/infrastructure-engineer.md +31 -0
  213. package/plugins/just-vibe/references/profiles/integration-architect.md +31 -0
  214. package/plugins/just-vibe/references/profiles/integration-engineer.md +31 -0
  215. package/plugins/just-vibe/references/profiles/ios-engineer.md +31 -0
  216. package/plugins/just-vibe/references/profiles/iot-engineer.md +31 -0
  217. package/plugins/just-vibe/references/profiles/kubernetes-engineer.md +31 -0
  218. package/plugins/just-vibe/references/profiles/llm-engineer.md +31 -0
  219. package/plugins/just-vibe/references/profiles/machine-learning-engineer.md +31 -0
  220. package/plugins/just-vibe/references/profiles/ml-architect.md +31 -0
  221. package/plugins/just-vibe/references/profiles/ml-data-engineer.md +31 -0
  222. package/plugins/just-vibe/references/profiles/ml-platform-engineer.md +31 -0
  223. package/plugins/just-vibe/references/profiles/mlops-engineer.md +31 -0
  224. package/plugins/just-vibe/references/profiles/mobile-engineer.md +31 -0
  225. package/plugins/just-vibe/references/profiles/network-engineer.md +31 -0
  226. package/plugins/just-vibe/references/profiles/nlp-engineer.md +31 -0
  227. package/plugins/just-vibe/references/profiles/observability-engineer.md +31 -0
  228. package/plugins/just-vibe/references/profiles/performance-engineer.md +31 -0
  229. package/plugins/just-vibe/references/profiles/platform-architect.md +31 -0
  230. package/plugins/just-vibe/references/profiles/platform-engineer.md +31 -0
  231. package/plugins/just-vibe/references/profiles/principal-engineer.md +31 -0
  232. package/plugins/just-vibe/references/profiles/privacy-engineer.md +31 -0
  233. package/plugins/just-vibe/references/profiles/product-engineer.md +31 -0
  234. package/plugins/just-vibe/references/profiles/product-security-engineer.md +31 -0
  235. package/plugins/just-vibe/references/profiles/protocol-engineer.md +31 -0
  236. package/plugins/just-vibe/references/profiles/qa-automation-engineer.md +31 -0
  237. package/plugins/just-vibe/references/profiles/recommendation-engineer.md +31 -0
  238. package/plugins/just-vibe/references/profiles/reinforcement-learning-engineer.md +31 -0
  239. package/plugins/just-vibe/references/profiles/research-engineer.md +31 -0
  240. package/plugins/just-vibe/references/profiles/research-scientist.md +31 -0
  241. package/plugins/just-vibe/references/profiles/responsible-ai-engineer.md +31 -0
  242. package/plugins/just-vibe/references/profiles/robotics-engineer.md +31 -0
  243. package/plugins/just-vibe/references/profiles/runtime-engineer.md +31 -0
  244. package/plugins/just-vibe/references/profiles/scientific-software-engineer.md +31 -0
  245. package/plugins/just-vibe/references/profiles/search-engineer.md +31 -0
  246. package/plugins/just-vibe/references/profiles/security-architect.md +31 -0
  247. package/plugins/just-vibe/references/profiles/security-automation-engineer.md +31 -0
  248. package/plugins/just-vibe/references/profiles/security-incident-responder.md +31 -0
  249. package/plugins/just-vibe/references/profiles/senior-software-engineer.md +31 -0
  250. package/plugins/just-vibe/references/profiles/simulation-engineer.md +31 -0
  251. package/plugins/just-vibe/references/profiles/site-reliability-engineer.md +31 -0
  252. package/plugins/just-vibe/references/profiles/software-architect.md +31 -0
  253. package/plugins/just-vibe/references/profiles/solutions-architect.md +31 -0
  254. package/plugins/just-vibe/references/profiles/speech-engineer.md +31 -0
  255. package/plugins/just-vibe/references/profiles/staff-engineer.md +31 -0
  256. package/plugins/just-vibe/references/profiles/storage-engineer.md +31 -0
  257. package/plugins/just-vibe/references/profiles/streaming-data-engineer.md +31 -0
  258. package/plugins/just-vibe/references/profiles/supply-chain-security-engineer.md +31 -0
  259. package/plugins/just-vibe/references/profiles/systems-engineer.md +31 -0
  260. package/plugins/just-vibe/references/profiles/tech-lead.md +31 -0
  261. package/plugins/just-vibe/references/profiles/technical-writer.md +31 -0
  262. package/plugins/just-vibe/references/profiles/test-infrastructure-engineer.md +31 -0
  263. package/plugins/just-vibe/references/profiles/ui-engineer.md +31 -0
  264. package/plugins/just-vibe/references/profiles/ux-engineer.md +31 -0
  265. package/plugins/just-vibe/references/profiles/web-performance-engineer.md +31 -0
  266. package/plugins/just-vibe/references/profiles/xr-engineer.md +31 -0
  267. package/plugins/just-vibe/references/profiles.md +59 -0
  268. package/plugins/just-vibe/references/runtime.md +70 -0
  269. package/plugins/just-vibe/references/scenarios/auth.md +31 -0
  270. package/plugins/just-vibe/references/scenarios/combobox.md +9 -0
  271. package/plugins/just-vibe/references/scenarios/date-picker.md +9 -0
  272. package/plugins/just-vibe/references/scenarios/delivery-evidence.md +21 -0
  273. package/plugins/just-vibe/references/scenarios/dialog.md +9 -0
  274. package/plugins/just-vibe/references/scenarios/training.md +21 -0
  275. package/plugins/just-vibe/references/teach-test.md +37 -0
  276. package/plugins/just-vibe/references/teaching.md +34 -0
  277. package/plugins/just-vibe/references/validation.md +11 -0
  278. package/plugins/just-vibe/scripts/discover-capabilities.mjs +3 -0
  279. package/plugins/just-vibe/scripts/hooks.mjs +14 -0
  280. package/plugins/just-vibe/scripts/inspect-project.mjs +3 -0
  281. package/plugins/just-vibe/scripts/installer.mjs +280 -0
  282. package/plugins/just-vibe/scripts/lib/automation.mjs +142 -0
  283. package/plugins/just-vibe/scripts/lib/bundle.mjs +100 -0
  284. package/plugins/just-vibe/scripts/lib/catalog.mjs +135 -0
  285. package/plugins/just-vibe/scripts/lib/command.mjs +26 -0
  286. package/plugins/just-vibe/scripts/lib/continuity.mjs +77 -0
  287. package/plugins/just-vibe/scripts/lib/discovery.mjs +84 -0
  288. package/plugins/just-vibe/scripts/lib/entrypoint.mjs +12 -0
  289. package/plugins/just-vibe/scripts/lib/evidence.mjs +136 -0
  290. package/plugins/just-vibe/scripts/lib/process.mjs +44 -0
  291. package/plugins/just-vibe/scripts/lib/profiles.mjs +83 -0
  292. package/plugins/just-vibe/scripts/lib/project.mjs +60 -0
  293. package/plugins/just-vibe/scripts/lib/routing.mjs +82 -0
  294. package/plugins/just-vibe/scripts/lib/run.mjs +248 -0
  295. package/plugins/just-vibe/scripts/lib/storage.mjs +84 -0
  296. package/plugins/just-vibe/scripts/lib/teaching.mjs +118 -0
  297. package/plugins/just-vibe/scripts/toolkit.mjs +225 -0
  298. package/plugins/just-vibe/skills/a11y/SKILL.md +8 -0
  299. package/plugins/just-vibe/skills/api-breaking/SKILL.md +56 -0
  300. package/plugins/just-vibe/skills/api-client/SKILL.md +56 -0
  301. package/plugins/just-vibe/skills/api-contract-test/SKILL.md +56 -0
  302. package/plugins/just-vibe/skills/api-design/SKILL.md +56 -0
  303. package/plugins/just-vibe/skills/api-errors/SKILL.md +56 -0
  304. package/plugins/just-vibe/skills/api-openapi/SKILL.md +56 -0
  305. package/plugins/just-vibe/skills/api-pagination/SKILL.md +56 -0
  306. package/plugins/just-vibe/skills/api-webhooks/SKILL.md +56 -0
  307. package/plugins/just-vibe/skills/arch-boundaries/SKILL.md +58 -0
  308. package/plugins/just-vibe/skills/arch-contracts/SKILL.md +56 -0
  309. package/plugins/just-vibe/skills/arch-event-flow/SKILL.md +56 -0
  310. package/plugins/just-vibe/skills/arch-feature/SKILL.md +57 -0
  311. package/plugins/just-vibe/skills/arch-map/SKILL.md +56 -0
  312. package/plugins/just-vibe/skills/arch-modernize/SKILL.md +56 -0
  313. package/plugins/just-vibe/skills/arch-scale/SKILL.md +58 -0
  314. package/plugins/just-vibe/skills/arch-tenancy/SKILL.md +56 -0
  315. package/plugins/just-vibe/skills/auto/SKILL.md +67 -0
  316. package/plugins/just-vibe/skills/automate/SKILL.md +56 -0
  317. package/plugins/just-vibe/skills/backend-auth/SKILL.md +63 -0
  318. package/plugins/just-vibe/skills/backend-cache/SKILL.md +59 -0
  319. package/plugins/just-vibe/skills/backend-concurrency/SKILL.md +58 -0
  320. package/plugins/just-vibe/skills/backend-idempotency/SKILL.md +58 -0
  321. package/plugins/just-vibe/skills/backend-jobs/SKILL.md +56 -0
  322. package/plugins/just-vibe/skills/backend-permissions/SKILL.md +56 -0
  323. package/plugins/just-vibe/skills/backend-resilience/SKILL.md +56 -0
  324. package/plugins/just-vibe/skills/backend-service/SKILL.md +56 -0
  325. package/plugins/just-vibe/skills/brainstorm/SKILL.md +56 -0
  326. package/plugins/just-vibe/skills/build/SKILL.md +56 -0
  327. package/plugins/just-vibe/skills/challenge/SKILL.md +56 -0
  328. package/plugins/just-vibe/skills/checkpoint/SKILL.md +59 -0
  329. package/plugins/just-vibe/skills/ci/SKILL.md +56 -0
  330. package/plugins/just-vibe/skills/cleanup/SKILL.md +56 -0
  331. package/plugins/just-vibe/skills/compare/SKILL.md +56 -0
  332. package/plugins/just-vibe/skills/copy/SKILL.md +56 -0
  333. package/plugins/just-vibe/skills/coverage/SKILL.md +56 -0
  334. package/plugins/just-vibe/skills/data-backfill/SKILL.md +56 -0
  335. package/plugins/just-vibe/skills/data-contract/SKILL.md +56 -0
  336. package/plugins/just-vibe/skills/data-incremental/SKILL.md +56 -0
  337. package/plugins/just-vibe/skills/data-lineage/SKILL.md +56 -0
  338. package/plugins/just-vibe/skills/data-pipeline/SKILL.md +56 -0
  339. package/plugins/just-vibe/skills/data-profile/SKILL.md +56 -0
  340. package/plugins/just-vibe/skills/data-quality/SKILL.md +56 -0
  341. package/plugins/just-vibe/skills/data-reconcile/SKILL.md +56 -0
  342. package/plugins/just-vibe/skills/db-access/SKILL.md +56 -0
  343. package/plugins/just-vibe/skills/db-explain/SKILL.md +56 -0
  344. package/plugins/just-vibe/skills/db-index/SKILL.md +56 -0
  345. package/plugins/just-vibe/skills/db-integrity/SKILL.md +57 -0
  346. package/plugins/just-vibe/skills/db-locks/SKILL.md +56 -0
  347. package/plugins/just-vibe/skills/db-migrate/SKILL.md +63 -0
  348. package/plugins/just-vibe/skills/db-query/SKILL.md +56 -0
  349. package/plugins/just-vibe/skills/db-schema/SKILL.md +56 -0
  350. package/plugins/just-vibe/skills/debug/SKILL.md +56 -0
  351. package/plugins/just-vibe/skills/decide/SKILL.md +57 -0
  352. package/plugins/just-vibe/skills/decision-adr/SKILL.md +56 -0
  353. package/plugins/just-vibe/skills/decision-buy-build/SKILL.md +56 -0
  354. package/plugins/just-vibe/skills/decision-matrix/SKILL.md +56 -0
  355. package/plugins/just-vibe/skills/decision-premortem/SKILL.md +56 -0
  356. package/plugins/just-vibe/skills/decision-reversible/SKILL.md +56 -0
  357. package/plugins/just-vibe/skills/decision-revisit/SKILL.md +56 -0
  358. package/plugins/just-vibe/skills/decision-spike/SKILL.md +58 -0
  359. package/plugins/just-vibe/skills/deploy/SKILL.md +56 -0
  360. package/plugins/just-vibe/skills/deps/SKILL.md +56 -0
  361. package/plugins/just-vibe/skills/design/SKILL.md +56 -0
  362. package/plugins/just-vibe/skills/do/SKILL.md +8 -0
  363. package/plugins/just-vibe/skills/docs/SKILL.md +56 -0
  364. package/plugins/just-vibe/skills/doctor/SKILL.md +58 -0
  365. package/plugins/just-vibe/skills/explain/SKILL.md +56 -0
  366. package/plugins/just-vibe/skills/fix/SKILL.md +57 -0
  367. package/plugins/just-vibe/skills/git-bisect/SKILL.md +58 -0
  368. package/plugins/just-vibe/skills/git-commit/SKILL.md +59 -0
  369. package/plugins/just-vibe/skills/git-conflicts/SKILL.md +58 -0
  370. package/plugins/just-vibe/skills/git-diff/SKILL.md +58 -0
  371. package/plugins/just-vibe/skills/git-recover/SKILL.md +58 -0
  372. package/plugins/just-vibe/skills/git-split/SKILL.md +58 -0
  373. package/plugins/just-vibe/skills/git-status/SKILL.md +58 -0
  374. package/plugins/just-vibe/skills/git-worktree/SKILL.md +58 -0
  375. package/plugins/just-vibe/skills/github-actions/SKILL.md +64 -0
  376. package/plugins/just-vibe/skills/github-address-review/SKILL.md +58 -0
  377. package/plugins/just-vibe/skills/github-fix-ci/SKILL.md +63 -0
  378. package/plugins/just-vibe/skills/github-issue/SKILL.md +58 -0
  379. package/plugins/just-vibe/skills/github-pr/SKILL.md +63 -0
  380. package/plugins/just-vibe/skills/github-release/SKILL.md +58 -0
  381. package/plugins/just-vibe/skills/github-review/SKILL.md +58 -0
  382. package/plugins/just-vibe/skills/github-triage/SKILL.md +58 -0
  383. package/plugins/just-vibe/skills/handoff/SKILL.md +62 -0
  384. package/plugins/just-vibe/skills/help/SKILL.md +63 -0
  385. package/plugins/just-vibe/skills/integrate/SKILL.md +56 -0
  386. package/plugins/just-vibe/skills/learn/SKILL.md +56 -0
  387. package/plugins/just-vibe/skills/llm-cost/SKILL.md +56 -0
  388. package/plugins/just-vibe/skills/llm-evals/SKILL.md +58 -0
  389. package/plugins/just-vibe/skills/llm-injection/SKILL.md +56 -0
  390. package/plugins/just-vibe/skills/llm-prompt/SKILL.md +56 -0
  391. package/plugins/just-vibe/skills/llm-rag/SKILL.md +56 -0
  392. package/plugins/just-vibe/skills/llm-retrieval/SKILL.md +56 -0
  393. package/plugins/just-vibe/skills/llm-structured/SKILL.md +56 -0
  394. package/plugins/just-vibe/skills/llm-tools/SKILL.md +58 -0
  395. package/plugins/just-vibe/skills/map/SKILL.md +56 -0
  396. package/plugins/just-vibe/skills/match/SKILL.md +56 -0
  397. package/plugins/just-vibe/skills/migrate/SKILL.md +56 -0
  398. package/plugins/just-vibe/skills/ml-ablation/SKILL.md +56 -0
  399. package/plugins/just-vibe/skills/ml-baseline/SKILL.md +56 -0
  400. package/plugins/just-vibe/skills/ml-batch/SKILL.md +56 -0
  401. package/plugins/just-vibe/skills/ml-calibrate/SKILL.md +56 -0
  402. package/plugins/just-vibe/skills/ml-dataset/SKILL.md +56 -0
  403. package/plugins/just-vibe/skills/ml-dataset-version/SKILL.md +56 -0
  404. package/plugins/just-vibe/skills/ml-debug-training/SKILL.md +56 -0
  405. package/plugins/just-vibe/skills/ml-drift/SKILL.md +56 -0
  406. package/plugins/just-vibe/skills/ml-error-analysis/SKILL.md +56 -0
  407. package/plugins/just-vibe/skills/ml-evaluate/SKILL.md +56 -0
  408. package/plugins/just-vibe/skills/ml-experiments/SKILL.md +56 -0
  409. package/plugins/just-vibe/skills/ml-explain/SKILL.md +56 -0
  410. package/plugins/just-vibe/skills/ml-features/SKILL.md +59 -0
  411. package/plugins/just-vibe/skills/ml-frame/SKILL.md +56 -0
  412. package/plugins/just-vibe/skills/ml-imbalance/SKILL.md +56 -0
  413. package/plugins/just-vibe/skills/ml-inference-perf/SKILL.md +56 -0
  414. package/plugins/just-vibe/skills/ml-labels/SKILL.md +56 -0
  415. package/plugins/just-vibe/skills/ml-leakage/SKILL.md +64 -0
  416. package/plugins/just-vibe/skills/ml-monitor/SKILL.md +56 -0
  417. package/plugins/just-vibe/skills/ml-package/SKILL.md +56 -0
  418. package/plugins/just-vibe/skills/ml-parity/SKILL.md +56 -0
  419. package/plugins/just-vibe/skills/ml-report/SKILL.md +56 -0
  420. package/plugins/just-vibe/skills/ml-reproduce/SKILL.md +56 -0
  421. package/plugins/just-vibe/skills/ml-robustness/SKILL.md +56 -0
  422. package/plugins/just-vibe/skills/ml-rollout/SKILL.md +56 -0
  423. package/plugins/just-vibe/skills/ml-serving/SKILL.md +56 -0
  424. package/plugins/just-vibe/skills/ml-slices/SKILL.md +56 -0
  425. package/plugins/just-vibe/skills/ml-split/SKILL.md +58 -0
  426. package/plugins/just-vibe/skills/ml-threshold/SKILL.md +56 -0
  427. package/plugins/just-vibe/skills/ml-train/SKILL.md +62 -0
  428. package/plugins/just-vibe/skills/ml-training-cost/SKILL.md +56 -0
  429. package/plugins/just-vibe/skills/ml-tune/SKILL.md +56 -0
  430. package/plugins/just-vibe/skills/ops-alerts/SKILL.md +56 -0
  431. package/plugins/just-vibe/skills/ops-container/SKILL.md +56 -0
  432. package/plugins/just-vibe/skills/ops-incident/SKILL.md +56 -0
  433. package/plugins/just-vibe/skills/ops-logs/SKILL.md +56 -0
  434. package/plugins/just-vibe/skills/ops-observability/SKILL.md +56 -0
  435. package/plugins/just-vibe/skills/ops-postmortem/SKILL.md +56 -0
  436. package/plugins/just-vibe/skills/ops-restore/SKILL.md +56 -0
  437. package/plugins/just-vibe/skills/ops-runbook/SKILL.md +56 -0
  438. package/plugins/just-vibe/skills/orient/SKILL.md +58 -0
  439. package/plugins/just-vibe/skills/perf/SKILL.md +56 -0
  440. package/plugins/just-vibe/skills/plan/SKILL.md +56 -0
  441. package/plugins/just-vibe/skills/polish/SKILL.md +56 -0
  442. package/plugins/just-vibe/skills/pr/SKILL.md +58 -0
  443. package/plugins/just-vibe/skills/profile/SKILL.md +66 -0
  444. package/plugins/just-vibe/skills/profiles/SKILL.md +58 -0
  445. package/plugins/just-vibe/skills/react-async/SKILL.md +57 -0
  446. package/plugins/just-vibe/skills/react-audit/SKILL.md +56 -0
  447. package/plugins/just-vibe/skills/react-component/SKILL.md +65 -0
  448. package/plugins/just-vibe/skills/react-effects/SKILL.md +58 -0
  449. package/plugins/just-vibe/skills/react-forms/SKILL.md +56 -0
  450. package/plugins/just-vibe/skills/react-hydration/SKILL.md +57 -0
  451. package/plugins/just-vibe/skills/react-rerenders/SKILL.md +56 -0
  452. package/plugins/just-vibe/skills/react-state/SKILL.md +56 -0
  453. package/plugins/just-vibe/skills/refactor/SKILL.md +56 -0
  454. package/plugins/just-vibe/skills/release/SKILL.md +58 -0
  455. package/plugins/just-vibe/skills/remember/SKILL.md +70 -0
  456. package/plugins/just-vibe/skills/repro/SKILL.md +56 -0
  457. package/plugins/just-vibe/skills/research/SKILL.md +56 -0
  458. package/plugins/just-vibe/skills/responsive/SKILL.md +8 -0
  459. package/plugins/just-vibe/skills/resume/SKILL.md +63 -0
  460. package/plugins/just-vibe/skills/review/SKILL.md +56 -0
  461. package/plugins/just-vibe/skills/scope/SKILL.md +56 -0
  462. package/plugins/just-vibe/skills/security/SKILL.md +56 -0
  463. package/plugins/just-vibe/skills/security-authz/SKILL.md +56 -0
  464. package/plugins/just-vibe/skills/security-config/SKILL.md +56 -0
  465. package/plugins/just-vibe/skills/security-dependencies/SKILL.md +56 -0
  466. package/plugins/just-vibe/skills/security-fix/SKILL.md +56 -0
  467. package/plugins/just-vibe/skills/security-inputs/SKILL.md +56 -0
  468. package/plugins/just-vibe/skills/security-secrets/SKILL.md +56 -0
  469. package/plugins/just-vibe/skills/security-threat-model/SKILL.md +56 -0
  470. package/plugins/just-vibe/skills/security-uploads/SKILL.md +56 -0
  471. package/plugins/just-vibe/skills/setup/SKILL.md +60 -0
  472. package/plugins/just-vibe/skills/skill/SKILL.md +56 -0
  473. package/plugins/just-vibe/skills/spec/SKILL.md +56 -0
  474. package/plugins/just-vibe/skills/tasks/SKILL.md +56 -0
  475. package/plugins/just-vibe/skills/teach/SKILL.md +63 -0
  476. package/plugins/just-vibe/skills/teach-test/SKILL.md +65 -0
  477. package/plugins/just-vibe/skills/test/SKILL.md +56 -0
  478. package/plugins/just-vibe/skills/test-e2e/SKILL.md +56 -0
  479. package/plugins/just-vibe/skills/test-fixtures/SKILL.md +56 -0
  480. package/plugins/just-vibe/skills/test-flaky/SKILL.md +56 -0
  481. package/plugins/just-vibe/skills/test-integration/SKILL.md +56 -0
  482. package/plugins/just-vibe/skills/test-load/SKILL.md +56 -0
  483. package/plugins/just-vibe/skills/test-property/SKILL.md +56 -0
  484. package/plugins/just-vibe/skills/test-regression/SKILL.md +57 -0
  485. package/plugins/just-vibe/skills/test-unit/SKILL.md +56 -0
  486. package/plugins/just-vibe/skills/tools/SKILL.md +64 -0
  487. package/plugins/just-vibe/skills/trace/SKILL.md +56 -0
  488. package/plugins/just-vibe/skills/ui-accessibility/SKILL.md +57 -0
  489. package/plugins/just-vibe/skills/ui-audit/SKILL.md +56 -0
  490. package/plugins/just-vibe/skills/ui-flow/SKILL.md +56 -0
  491. package/plugins/just-vibe/skills/ui-motion/SKILL.md +56 -0
  492. package/plugins/just-vibe/skills/ui-responsive/SKILL.md +56 -0
  493. package/plugins/just-vibe/skills/ui-states/SKILL.md +56 -0
  494. package/plugins/just-vibe/skills/ui-system/SKILL.md +56 -0
  495. package/plugins/just-vibe/skills/ui-visual-diff/SKILL.md +56 -0
  496. package/plugins/just-vibe/skills/vercel-audit/SKILL.md +56 -0
  497. package/plugins/just-vibe/skills/vercel-build-fix/SKILL.md +63 -0
  498. package/plugins/just-vibe/skills/vercel-env/SKILL.md +56 -0
  499. package/plugins/just-vibe/skills/vercel-performance/SKILL.md +56 -0
  500. package/plugins/just-vibe/skills/vercel-preview/SKILL.md +56 -0
  501. package/plugins/just-vibe/skills/vercel-release-check/SKILL.md +57 -0
  502. package/plugins/just-vibe/skills/vercel-routing/SKILL.md +56 -0
  503. package/plugins/just-vibe/skills/vercel-runtime/SKILL.md +61 -0
  504. package/plugins/just-vibe/skills/verify/SKILL.md +59 -0
  505. package/plugins/just-vibe/skills/vite-assets/SKILL.md +57 -0
  506. package/plugins/just-vibe/skills/vite-bundle/SKILL.md +58 -0
  507. package/plugins/just-vibe/skills/vite-chunks/SKILL.md +56 -0
  508. package/plugins/just-vibe/skills/vite-config/SKILL.md +56 -0
  509. package/plugins/just-vibe/skills/vite-env/SKILL.md +56 -0
  510. package/plugins/just-vibe/skills/vite-hmr/SKILL.md +56 -0
  511. package/plugins/just-vibe/skills/vite-setup/SKILL.md +56 -0
  512. package/plugins/just-vibe/skills/vite-upgrade/SKILL.md +56 -0
@@ -0,0 +1,56 @@
1
+ ---
2
+ name: ml-drift
3
+ description: "Design checks for input or prediction-distribution changes Use for observed input/prediction distribution change; ml-evaluate requires outcomes to establish quality."
4
+ ---
5
+
6
+ # ml-drift
7
+
8
+ Design checks for input or prediction-distribution changes
9
+
10
+ ## Choose this workflow
11
+
12
+ Use for observed input/prediction distribution change; ml-evaluate requires outcomes to establish quality.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML deployment methods](../../references/packs/ml-deployment.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **plan**. Plan; reference/current data or prediction windows, features, seasonality, and sensitivity requirements.
19
+
20
+ versioned model and preprocessing artifacts, input/output schema, runtime/dependencies, operating targets, and authorized environment. Validate artifact trust before loading formats that can execute code. Packaging or writing monitoring configuration does not deploy a model or enable a hosted service.
21
+
22
+ Declared evidence requirements: `ml.artifacts`. Use actual host discovery or adequate supplied artifacts; unavailable evidence remains blocked/unknown.
23
+
24
+ ## Scope
25
+
26
+ Distribution-change detection and investigation; no automatic retraining.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Align schemas/windows, choose meaningful per-feature and aggregate checks, account for sample size/seasonality, inspect effect sizes, and define follow-up on signals.
33
+ - Align schema, sampling and seasonal windows, compare effect sizes and support changes and separate data-pipeline changes from population changes.
34
+
35
+ ## Decision branches
36
+
37
+ - **When drift is statistically significant but outcomes are unavailable:** Report a monitoring signal and follow-up, not confirmed accuracy degradation.
38
+
39
+ ## Deliver and verify
40
+
41
+ - Drift protocol or report with baselines, thresholds, uncertainty, and investigation guidance.
42
+ - Baseline/current identities, shift measures, sample sizes and next evidence needed.
43
+
44
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
45
+
46
+ - A changed categorical support is detected; a statistically significant tiny shift is not automatically called model failure.
47
+
48
+ ## Stop and recover
49
+
50
+ - Drift is not proof of accuracy degradation without outcome evidence. Baseline refresh requires explicit policy, not silent adaptation.
51
+
52
+ ## Example requests
53
+
54
+ - **Normal (plan):** Plan distribution checks that account for seasonality and sample size.
55
+ - **edge (plan):** Investigate a new categorical value and a seasonal traffic shift.
56
+ - **blocked (inspect):** Assess drift from aggregates without inferring model failure or silently refreshing the baseline.
@@ -0,0 +1,56 @@
1
+ ---
2
+ name: ml-error-analysis
3
+ description: "Group failures into actionable patterns and examples Use to inspect model mistakes; ml-slices computes cohort metrics and ml-debug-training diagnoses optimization."
4
+ ---
5
+
6
+ # ml-error-analysis
7
+
8
+ Group failures into actionable patterns and examples
9
+
10
+ ## Choose this workflow
11
+
12
+ Use to inspect model mistakes; ml-slices computes cohort metrics and ml-debug-training diagnoses optimization.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML evaluation methods](../../references/packs/ml-evaluation.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **inspect**. Inspect; predictions, labels, task costs, and permitted redacted examples.
19
+
20
+ frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
21
+
22
+ Declared evidence requirements: `ml.artifacts`. Use actual host discovery or adequate supplied artifacts; unavailable evidence remains blocked/unknown.
23
+
24
+ ## Scope
25
+
26
+ Actionable failure patterns; no automatic retraining or relabeling.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Define errors according to task, group by meaningful factors, inspect representative cases and denominators, distinguish label problems, and propose targeted next experiments.
33
+ - Define the error event and denominator, group by meaningful factors and compare representative failures with matched successes and possible label problems.
34
+
35
+ ## Decision branches
36
+
37
+ - **When the analysis uses held-out test outcomes:** Keep findings exploratory and require fresh confirmation before tuning to those patterns.
38
+
39
+ ## Deliver and verify
40
+
41
+ - Error taxonomy, frequency/impact evidence, examples, and interventions to test.
42
+ - Error taxonomy, cohort counts, examples and proposed discriminating experiments.
43
+
44
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
45
+
46
+ - Common groups are judged against their population size; a few vivid cases do not imply prevalence.
47
+
48
+ ## Stop and recover
49
+
50
+ - Protect sensitive records. Patterns discovered on test data must not become tuning targets without a fresh evaluation plan.
51
+
52
+ ## Example requests
53
+
54
+ - **Normal (inspect):** Group failures from these predictions into actionable patterns with denominators.
55
+ - **edge (inspect):** Analyze rare high-cost errors without treating vivid examples as prevalence.
56
+ - **blocked (inspect):** Analyze aggregate errors when sensitive examples cannot be inspected.
@@ -0,0 +1,56 @@
1
+ ---
2
+ name: ml-evaluate
3
+ description: "Evaluate using task-appropriate metrics and baselines Use for fixed-model evaluation; ml-threshold and ml-calibrate require separate selection data."
4
+ ---
5
+
6
+ # ml-evaluate
7
+
8
+ Evaluate using task-appropriate metrics and baselines
9
+
10
+ ## Choose this workflow
11
+
12
+ Use for fixed-model evaluation; ml-threshold and ml-calibrate require separate selection data.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML evaluation methods](../../references/packs/ml-evaluation.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **plan**. Plan; model, dataset/split, task metrics, baseline, and inference budget. Explicit evaluation execution selects apply.
19
+
20
+ frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
21
+
22
+ Declared evidence requirements: `ml.artifacts`. Use actual host discovery or adequate supplied artifacts; unavailable evidence remains blocked/unknown.
23
+
24
+ ## Scope
25
+
26
+ Fixed-protocol performance evaluation, not model tuning.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Validate alignment and eligibility, run authorized predictions, compute declared metrics and appropriate uncertainty, compare baseline, and record excluded/missing cases.
33
+ - Align predictions and labels by stable row identity, freeze eligibility/metric definitions and count missing, excluded and failed predictions before computing results.
34
+
35
+ ## Decision branches
36
+
37
+ - **When observations are dependent within entities or time blocks:** Use an uncertainty method matching that dependence or explicitly leave uncertainty unestimated.
38
+
39
+ ## Deliver and verify
40
+
41
+ - Reproducible evaluation report with data/model identity, metrics, denominators, and limitations.
42
+ - Model/data/split identity, denominators, baseline metrics and uncertainty method.
43
+
44
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
45
+
46
+ - Prediction/label misalignment fails validation; missing predictions are counted rather than silently dropped into better scores.
47
+
48
+ ## Stop and recover
49
+
50
+ - No test-set-driven changes during evaluation. Missing ground truth restricts output to operational/descriptive checks.
51
+
52
+ ## Example requests
53
+
54
+ - **Normal (plan):** Plan evaluating the frozen model against the declared baseline and untouched test split.
55
+ - **edge (plan):** Evaluate shuffled prediction rows with missing outputs and delayed labels.
56
+ - **blocked (inspect):** Evaluate operational behavior without ground truth; do not report accuracy.
@@ -0,0 +1,56 @@
1
+ ---
2
+ name: ml-experiments
3
+ description: "Compare runs and check that data and evaluation conditions match Use to compare recorded runs; ml-train produces a run and ml-report communicates validated conclusions."
4
+ ---
5
+
6
+ # ml-experiments
7
+
8
+ Compare runs and check that data and evaluation conditions match
9
+
10
+ ## Choose this workflow
11
+
12
+ Use to compare recorded runs; ml-train produces a run and ml-report communicates validated conclusions.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML experimentation methods](../../references/packs/ml-experiments.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **inspect**. Inspect; run IDs/artifacts, metric of interest, and comparison scope.
19
+
20
+ dataset/split manifests, fixed objective/metric, environment/dependencies, baseline where applicable, and explicit compute limits. Record code revision, configuration, seeds, artifact paths, and resource use. Local smoke checks do not imply authorization for paid training. Never optimize on the held-out test set.
21
+
22
+ Declared evidence requirements: `ml.artifacts`. Use actual host discovery or adequate supplied artifacts; unavailable evidence remains blocked/unknown.
23
+
24
+ ## Scope
25
+
26
+ Compare existing experiments and their compatibility; no automatic reruns.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Reconcile data/split/code/config identities, normalize metric definitions, inspect failed/missing runs, compare quality and resources, and separate incompatible cohorts.
33
+ - Reconcile dataset/split/code/config identities, metric denominator and selection history; include failed/pruned runs in total resource accounting.
34
+
35
+ ## Decision branches
36
+
37
+ - **When runs used different populations or metric definitions:** Group them separately and propose a matched comparison instead of a misleading leaderboard.
38
+
39
+ ## Deliver and verify
40
+
41
+ - Experiment comparison, strongest supported result, and comparability gaps.
42
+ - Comparable-run groups, quality/resource evidence and missing metadata.
43
+
44
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
45
+
46
+ - Different test populations are not ranked as directly comparable; failed trials are not silently omitted from cost accounting.
47
+
48
+ ## Stop and recover
49
+
50
+ - Missing metadata prevents strong conclusions. Do not select a winner solely from rounded headline scores.
51
+
52
+ ## Example requests
53
+
54
+ - **Normal (inspect):** Compare these runs and flag different datasets or metric definitions.
55
+ - **edge (inspect):** Compare experiments that used different test periods and rounded headline scores.
56
+ - **blocked (inspect):** Inspect incomplete run exports without inventing missing metrics or costs.
@@ -0,0 +1,56 @@
1
+ ---
2
+ name: ml-explain
3
+ description: "Investigate behavior with appropriate explanation methods and limits Use to interpret model behavior; explain describes code and teach explains concepts."
4
+ ---
5
+
6
+ # ml-explain
7
+
8
+ Investigate behavior with appropriate explanation methods and limits
9
+
10
+ ## Choose this workflow
11
+
12
+ Use to interpret model behavior; explain describes code and teach explains concepts.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML evaluation methods](../../references/packs/ml-evaluation.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **inspect**. Inspect; model, prediction/global behavior question, data access, and audience.
19
+
20
+ frozen model/artifact, evaluation dataset identity, labels where needed, metric definitions, and task/operating context. Report sample counts and uncertainty appropriate to dependencies; avoid repeated test-set tuning. Exploratory findings need fresh confirmation before strong generalization claims.
21
+
22
+ Declared evidence requirements: `ml.artifacts`. Use actual host discovery or adequate supplied artifacts; unavailable evidence remains blocked/unknown.
23
+
24
+ ## Scope
25
+
26
+ Appropriate feature/behavior explanation with method limitations; not causal attribution by default.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Choose a method compatible with the model/question, inspect baseline/background dependence, check stability/correlated features, and connect explanations to actual examples.
33
+ - State whether the question concerns one prediction or global behavior, select a compatible method and examine background data and correlated-feature sensitivity.
34
+
35
+ ## Decision branches
36
+
37
+ - **When explanations vary strongly with reasonable baselines:** Report that dependence and avoid a causal or uniquely determined attribution claim.
38
+
39
+ ## Deliver and verify
40
+
41
+ - Explanation, method/settings, supporting evidence, and limits.
42
+ - Method/question fit, example explanations, stability checks and limitations.
43
+
44
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
45
+
46
+ - Correlated features are not treated as independent causal effects; unstable explanations are disclosed.
47
+
48
+ ## Stop and recover
49
+
50
+ - Expensive explanation runs need bounded execution. Do not expose proprietary or personal data through unnecessary examples.
51
+
52
+ ## Example requests
53
+
54
+ - **Normal (inspect):** Explain these predictions and separate feature association from causation.
55
+ - **edge (inspect):** Explain correlated feature importance without implying causation.
56
+ - **blocked (inspect):** Plan explanations without model artifacts or expensive inference authorization.
@@ -0,0 +1,59 @@
1
+ ---
2
+ name: ml-features
3
+ description: "Design features available at prediction time and test usefulness Use to implement prediction-time feature transformations; ml-labels defines outcomes and ml-leakage audits leakage."
4
+ ---
5
+
6
+ # ml-features
7
+
8
+ Design features available at prediction time and test usefulness
9
+
10
+ ## Choose this workflow
11
+
12
+ Use to implement prediction-time feature transformations; ml-labels defines outcomes and ml-leakage audits leakage.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML data methods](../../references/packs/ml-data.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **plan**. Plan; task, feature sources, availability timing, baseline, and evaluation protocol.
19
+
20
+ task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.
21
+
22
+ Declared evidence requirements: `data.read`. Use actual host discovery or adequate supplied artifacts; unavailable evidence remains blocked/unknown.
23
+
24
+ ## Scope
25
+
26
+ Predictable, reproducible feature engineering and controlled usefulness checks; implement/run when requested.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Define row grain, entity/event identity, prediction instant, feature event time, source availability time, revision order and missing behavior. Specify timezone and exact window endpoints. Missing availability evidence blocks historical-validity claims.
33
+ - For versioned sources, reconstruct the latest version actually available at each prediction instant before applying its event-time window. A later correction can move an event outside the window; filtering versions first can resurrect an obsolete value. Deduplicate by the documented event/version identity, scoped to the entity where required.
34
+ - Fit learned transforms only on eligible training rows within each simulated fit or cross-validation fold. Preserve input ordering/identity and define empty, constant, missing and unseen-category behavior; use the same transformation semantics at serving time.
35
+ - Verify exact time boundaries, mixed explicit offsets, late arrivals, corrections, duplicates, negative/zero values and no input mutation. Compare engineered features to an independent tiny example before proposing a usefulness experiment.
36
+ - Only claim usefulness after a controlled baseline comparison on the chosen evaluation protocol and compute budget. A correct feature builder alone establishes neither predictive gain nor production readiness.
37
+
38
+ ## Decision branches
39
+
40
+ - **When records can be corrected after their original event time:** Use availability and revision semantics to reconstruct the visible version first; never join historical predictions to today’s final mutable table without qualification.
41
+ - **When no mature training rows remain after eligibility filtering:** Use the documented empty-transform behavior or report the missing training prerequisite. Do not fit preprocessing on validation/test rows to avoid the empty case.
42
+
43
+ ## Deliver and verify
44
+
45
+ - Feature/time/identity contract, implementation when requested, eligible fitting population, boundary and revision evidence, and separately measured usefulness results if any.
46
+
47
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
48
+
49
+ - Historical features use only versions available at prediction time. Training transformations exclude held-out and ineligible rows, and preserve the documented empty/constant/missing behavior.
50
+
51
+ ## Stop and recover
52
+
53
+ - No test-set-driven feature selection. Unsupported availability timing blocks claims that a feature is deployable.
54
+
55
+ ## Example requests
56
+
57
+ - **Normal (plan):** Plan point-in-time customer features fitted only on training data.
58
+ - **edge (apply):** Build categorical features with unseen values and delayed source updates.
59
+ - **blocked (inspect):** Design features without trustworthy availability timestamps; do not claim deployability.
@@ -0,0 +1,56 @@
1
+ ---
2
+ name: ml-frame
3
+ description: "Define target, prediction moment, unit of analysis, and objective Use to define the prediction problem; ml-baseline implements the first comparator after the task is defined."
4
+ ---
5
+
6
+ # ml-frame
7
+
8
+ Define target, prediction moment, unit of analysis, and objective
9
+
10
+ ## Choose this workflow
11
+
12
+ Use to define the prediction problem; ml-baseline implements the first comparator after the task is defined.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML data methods](../../references/packs/ml-data.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **plan**. Plan; decision to support, population, available data, prediction timing, and operational objective.
19
+
20
+ task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.
21
+
22
+ Resolve any task-specific tools, target identity and evidence before dependent actions. No external connection is assumed.
23
+
24
+ ## Scope
25
+
26
+ Define the modeling problem before selecting algorithms.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Specify unit of analysis, target/label horizon, information available at prediction time, action taken from predictions, baseline, and costs of errors.
33
+ - State one prediction row's entity, timestamp, available information, label horizon and downstream action; compare a rule-based decision before choosing ML.
34
+
35
+ ## Decision branches
36
+
37
+ - **When label timing or intervention changes the observed outcome:** Separate prediction from causal/intervention claims and identify the missing observation process.
38
+
39
+ ## Deliver and verify
40
+
41
+ - Modeling brief with success metrics, eligibility/exclusions, deployment assumptions, and unresolved policy choices.
42
+ - Task card with row grain, prediction moment, outcome horizon, action and error costs.
43
+
44
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
45
+
46
+ - Target and prediction time are unambiguous; a proxy label's mismatch with the real objective is explicit.
47
+
48
+ ## Stop and recover
49
+
50
+ - Do not force an ML solution when deterministic rules suffice or invent business error costs without input.
51
+
52
+ ## Example requests
53
+
54
+ - **Normal (plan):** Frame churn prediction 30 days before cancellation, including unit and label horizon.
55
+ - **edge (plan):** Frame failure prediction for machines with delayed maintenance labels.
56
+ - **blocked (inspect):** Define the task without business error costs; keep threshold selection undecided.
@@ -0,0 +1,56 @@
1
+ ---
2
+ name: ml-imbalance
3
+ description: "Evaluate sampling, weighting, metrics, and thresholds for rare outcomes Use when rare outcomes affect metrics or training; ml-threshold selects operational decisions."
4
+ ---
5
+
6
+ # ml-imbalance
7
+
8
+ Evaluate sampling, weighting, metrics, and thresholds for rare outcomes
9
+
10
+ ## Choose this workflow
11
+
12
+ Use when rare outcomes affect metrics or training; ml-threshold selects operational decisions.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML data methods](../../references/packs/ml-data.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **plan**. Plan; prevalence, class definitions, error costs/capacity, and split protocol.
19
+
20
+ task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.
21
+
22
+ Declared evidence requirements: `data.read`. Use actual host discovery or adequate supplied artifacts; unavailable evidence remains blocked/unknown.
23
+
24
+ ## Scope
25
+
26
+ Sampling/weighting, evaluation, and operating-point options for rare outcomes.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Establish naive baselines, inspect per-class/sample counts, choose suitable metrics, compare resampling/weighting only within training folds, and assess deployment prevalence effects.
33
+ - Compute baseline prevalence and class counts by split, choose task-relevant precision/recall measures and restrict resampling to training folds.
34
+
35
+ ## Decision branches
36
+
37
+ - **When prevalence differs between sampled training and deployment:** Separate learned ranking from probability calibration and expected operational workload.
38
+
39
+ ## Deliver and verify
40
+
41
+ - Imbalance strategy with bounded experiments and threshold considerations.
42
+ - Baselines, per-class denominators, resampling protocol and uncertainty limits.
43
+
44
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
45
+
46
+ - A high-accuracy all-negative model is not accepted as useful; resampled training prevalence is not confused with deployed probability calibration.
47
+
48
+ ## Stop and recover
49
+
50
+ - Do not synthesize across validation/test boundaries or invent business tradeoffs. Low positive counts require uncertainty disclosure.
51
+
52
+ ## Example requests
53
+
54
+ - **Normal (plan):** Compare weighting and metrics for rare fraud with limited review capacity.
55
+ - **edge (plan):** Compare a high-accuracy all-negative baseline against a rare-event model.
56
+ - **blocked (inspect):** Assess imbalance with few positives; do not invent stable confidence or business costs.
@@ -0,0 +1,56 @@
1
+ ---
2
+ name: ml-inference-perf
3
+ description: "Measure latency, throughput, memory, and optimization tradeoffs Use for latency/throughput/resource benchmarking; ml-training-cost covers training."
4
+ ---
5
+
6
+ # ml-inference-perf
7
+
8
+ Measure latency, throughput, memory, and optimization tradeoffs
9
+
10
+ ## Choose this workflow
11
+
12
+ Use for latency/throughput/resource benchmarking; ml-training-cost covers training.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML deployment methods](../../references/packs/ml-deployment.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **plan**. Plan; model/runtime, input distribution, concurrency, hardware, quality floor, and benchmark budget.
19
+
20
+ versioned model and preprocessing artifacts, input/output schema, runtime/dependencies, operating targets, and authorized environment. Validate artifact trust before loading formats that can execute code. Packaging or writing monitoring configuration does not deploy a model or enable a hosted service.
21
+
22
+ Resolve any task-specific tools, target identity and evidence before dependent actions. No external connection is assumed.
23
+
24
+ ## Scope
25
+
26
+ Latency, throughput, memory, batching, warmup, and optimization tradeoffs.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Define comparable benchmark conditions, separate cold/warm paths, measure bounded authorized workloads, identify bottlenecks, and check quality after optimizations.
33
+ - Specify hardware, precision, batch/concurrency and payload distribution; separate load/warmup from steady-state and measure tail behavior within caps.
34
+
35
+ ## Decision branches
36
+
37
+ - **When quantization or batching improves speed:** Re-evaluate quality, memory and latency under the same workload before accepting it.
38
+
39
+ ## Deliver and verify
40
+
41
+ - Benchmark protocol/results and optimization recommendation or authorized patch.
42
+ - Benchmark conditions, sample size, cold/warm/tail metrics and quality comparison.
43
+
44
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
45
+
46
+ - Tail latency and throughput are reported under stated load; quantization gains include a quality comparison.
47
+
48
+ ## Stop and recover
49
+
50
+ - No production load or paid hardware implicitly. A single warm request cannot establish capacity.
51
+
52
+ ## Example requests
53
+
54
+ - **Normal (plan):** Plan a bounded benchmark for cold/warm latency and throughput at fixed quality.
55
+ - **edge (plan):** Benchmark batched inference under a latency deadline without hiding warmup cost.
56
+ - **blocked (inspect):** Plan a benchmark with no authorized hardware or production traffic budget.
@@ -0,0 +1,56 @@
1
+ ---
2
+ name: ml-labels
3
+ description: "Inspect label definitions, noise, disagreement, and missing outcomes Use for label construction and annotation quality; ml-leakage checks prediction-time information flow."
4
+ ---
5
+
6
+ # ml-labels
7
+
8
+ Inspect label definitions, noise, disagreement, and missing outcomes
9
+
10
+ ## Choose this workflow
11
+
12
+ Use for label construction and annotation quality; ml-leakage checks prediction-time information flow.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML data methods](../../references/packs/ml-data.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **inspect**. Inspect; label definitions, annotation/outcome sources, timing, and permitted samples.
19
+
20
+ task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.
21
+
22
+ Declared evidence requirements: `data.read`. Use actual host discovery or adequate supplied artifacts; unavailable evidence remains blocked/unknown.
23
+
24
+ ## Scope
25
+
26
+ Label consistency, noise, disagreement, censoring, and missing outcomes.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Trace label construction, compare annotations/outcomes, distinguish disagreement from ambiguous policy, inspect timing and coverage, and propose adjudication/quality checks.
33
+ - Trace label source, event horizon and maturity; distinguish true negatives, unobserved outcomes, contradictory annotations and policy ambiguity.
34
+
35
+ ## Decision branches
36
+
37
+ - **When annotators disagree on an ambiguous definition:** Preserve disagreement, clarify policy and adjudicate within scope rather than silently majority-voting it away.
38
+
39
+ ## Deliver and verify
40
+
41
+ - Label audit with concrete patterns, estimated rates with denominators, and corrective options.
42
+ - Label definition, maturity/coverage checks and reproducible disagreement examples.
43
+
44
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
45
+
46
+ - Unobserved outcomes are not automatically negative; conflicting annotations are tracked rather than silently overwritten.
47
+
48
+ ## Stop and recover
49
+
50
+ - Relabeling requires explicit policy and scope. Avoid exposing sensitive examples or claiming a single annotator is ground truth without justification.
51
+
52
+ ## Example requests
53
+
54
+ - **Normal (inspect):** Audit how missing outcome follow-up and annotation disagreement affect labels.
55
+ - **edge (inspect):** Audit labels when missing follow-up was encoded as no failure.
56
+ - **blocked (inspect):** Assess annotation policy without identifiable raw examples or relabeling permission.
@@ -0,0 +1,64 @@
1
+ ---
2
+ name: ml-leakage
3
+ description: "Find target leakage, temporal leakage, and split contamination Use to audit demonstrated information leakage; ml-split designs the evaluation protocol."
4
+ ---
5
+
6
+ # ml-leakage
7
+
8
+ Find target leakage, temporal leakage, and split contamination
9
+
10
+ ## Choose this workflow
11
+
12
+ Use to audit demonstrated information leakage; ml-split designs the evaluation protocol.
13
+
14
+ Read [shared execution](../../references/execution.md) for context/mode/authority handling and [ML data methods](../../references/packs/ml-data.md) for tool selection and operational details. Resolve these paths from this skill file; all runtime assets ship inside the plugin.
15
+
16
+ ## Input and mode
17
+
18
+ Use the complete request appended to this invocation, preserving all constraints and references. Default mode: **inspect**. Inspect; task/prediction moment, features, preprocessing, labels, and split lineage.
19
+
20
+ task definition, dataset identity, field semantics, entity/time keys, and permission to inspect bounded data. Record prediction moment, label horizon, sampling, and provenance. Preserve held-out evaluation boundaries; no data upload, label alteration, or feature fitting across splits implicitly.
21
+
22
+ Declared evidence requirements: `data.read`. Use actual host discovery or adequate supplied artifacts; unavailable evidence remains blocked/unknown.
23
+
24
+ ## Scope
25
+
26
+ Target proxies, future information, cross-split fitting, duplicates, and entity contamination.
27
+
28
+ None by default. Plan artifacts may be saved when requested.
29
+
30
+ ## Execute
31
+
32
+ - Identify the prediction moment, label horizon, feature availability, split membership and intended deployment population from supplied evidence. Mark missing definitions as unknown.
33
+ - Trace suspicious features to when their values could actually have been available. Distinguish future-derived values from valid point-in-time historical features.
34
+ - Inspect preprocessing fit/transform boundaries when code or fit records exist. Their absence means unverified, not proof of correct or incorrect fitting.
35
+ - For each finding record the exact observed rows or source paths, the conclusion those observations support, and any assumptions needed for a stronger conclusion. Separate confirmed defects, conditional risks and missing evidence in the report.
36
+ - Quantify entity, interval and outcome-horizon overlap. Do not infer identical raw measurements or a shared outcome event from metadata alone. Determine group separation from whether deployment targets known entities, new entities or new groups.
37
+ - Check label maturity against the simulated model-fit and prediction times. State the historical-deployment assumption when applying temporal cutoffs or an embargo; choose gaps from actual availability and overlap instead of a universal duration.
38
+ - Before delivering, check every claim labeled proven against its cited evidence. Correct unsupported absolutes, including assertions that all scores are invalid or a split is always wrong. Identify which scores would be affected under which assumptions, and require re-evaluation after confirmed leakage is corrected.
39
+ - Build a compact evidence ledger: field or row, availability time, prediction/fit time, observed violation, affected score and assumptions; keep overlap metadata separate from shared measurements or events.
40
+
41
+ ## Decision branches
42
+
43
+ - **When no raw measurements, event IDs or fitting history establish dependence:** Report conditional risk or unknown, not proven shared events, mandatory gap length or universal score invalidity.
44
+
45
+ ## Deliver and verify
46
+
47
+ - An evidence-backed leakage audit separating confirmed defects, conditional risks and unknowns; each finding names the supporting rows/source, assumptions, affected evaluation and correction or missing evidence.
48
+ - Finding ledger that ties every confirmed defect to supplied source or rows and bounds the affected evaluation.
49
+
50
+ Verify these observable conditions when applicable to the actual task; do not claim they were exercised from merely reading this file:
51
+
52
+ - Future-only features and immature training labels are flagged against the stated prediction/fit times; cross-split preprocessing fitting is detected only when source or fit history establishes it.
53
+ - Metadata-only overlap is quantified without asserting identical sensor values or a shared failure event. Group separation is conditional on the intended deployment population.
54
+ - Missing labels remain unknown outcomes, unavailable pipeline evidence remains unverified, and score invalidation is limited to affected evaluation assumptions rather than invented results.
55
+
56
+ ## Stop and recover
57
+
58
+ - Do not claim absence of leakage when provenance is missing. Remediation must invalidate affected scores rather than preserve misleading results.
59
+
60
+ ## Example requests
61
+
62
+ - **Normal (inspect):** Audit churn features for values unavailable 30 days before cancellation.
63
+ - **edge (inspect):** Audit overlapping windows whose metadata does not prove shared sensor values or outcome events.
64
+ - **blocked (inspect):** Review lineage with missing preprocessing code and event IDs; leave unsupported claims unknown.