@claude-flow/cli 3.32.9 → 3.32.11

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (447) hide show
  1. package/.claude/.proven-config-version +1 -0
  2. package/.claude/agents/analysis/analyze-code-quality.md +178 -178
  3. package/.claude/agents/analysis/code-analyzer.md +209 -209
  4. package/.claude/agents/analysis/code-review/analyze-code-quality.md +178 -178
  5. package/.claude/agents/architecture/arch-system-design.md +156 -156
  6. package/.claude/agents/architecture/system-design/arch-system-design.md +154 -154
  7. package/.claude/agents/browser/browser-agent.yaml +182 -182
  8. package/.claude/agents/consensus/byzantine-coordinator.md +62 -62
  9. package/.claude/agents/consensus/crdt-synchronizer.md +996 -996
  10. package/.claude/agents/consensus/gossip-coordinator.md +62 -62
  11. package/.claude/agents/consensus/performance-benchmarker.md +850 -850
  12. package/.claude/agents/consensus/quorum-manager.md +822 -822
  13. package/.claude/agents/consensus/raft-manager.md +62 -62
  14. package/.claude/agents/consensus/security-manager.md +621 -621
  15. package/.claude/agents/core/planner.md +374 -374
  16. package/.claude/agents/custom/test-long-runner.md +44 -44
  17. package/.claude/agents/data/data-ml-model.md +444 -444
  18. package/.claude/agents/data/ml/data-ml-model.md +192 -192
  19. package/.claude/agents/development/backend/dev-backend-api.md +141 -141
  20. package/.claude/agents/development/dev-backend-api.md +344 -344
  21. package/.claude/agents/devops/ci-cd/ops-cicd-github.md +163 -163
  22. package/.claude/agents/devops/ops-cicd-github.md +164 -164
  23. package/.claude/agents/documentation/api-docs/docs-api-openapi.md +173 -173
  24. package/.claude/agents/documentation/docs-api-openapi.md +354 -354
  25. package/.claude/agents/flow-nexus/app-store.md +87 -87
  26. package/.claude/agents/flow-nexus/authentication.md +68 -68
  27. package/.claude/agents/flow-nexus/challenges.md +80 -80
  28. package/.claude/agents/flow-nexus/neural-network.md +87 -87
  29. package/.claude/agents/flow-nexus/payments.md +82 -82
  30. package/.claude/agents/flow-nexus/sandbox.md +75 -75
  31. package/.claude/agents/flow-nexus/swarm.md +75 -75
  32. package/.claude/agents/flow-nexus/user-tools.md +95 -95
  33. package/.claude/agents/flow-nexus/workflow.md +83 -83
  34. package/.claude/agents/github/code-review-swarm.md +377 -377
  35. package/.claude/agents/github/github-modes.md +172 -172
  36. package/.claude/agents/github/issue-tracker.md +575 -575
  37. package/.claude/agents/github/multi-repo-swarm.md +552 -552
  38. package/.claude/agents/github/pr-manager.md +437 -437
  39. package/.claude/agents/github/project-board-sync.md +508 -508
  40. package/.claude/agents/github/release-manager.md +604 -604
  41. package/.claude/agents/github/release-swarm.md +582 -582
  42. package/.claude/agents/github/repo-architect.md +397 -397
  43. package/.claude/agents/github/swarm-issue.md +572 -572
  44. package/.claude/agents/github/swarm-pr.md +427 -427
  45. package/.claude/agents/github/sync-coordinator.md +451 -451
  46. package/.claude/agents/github/workflow-automation.md +902 -902
  47. package/.claude/agents/goal/agent.md +815 -815
  48. package/.claude/agents/optimization/benchmark-suite.md +664 -664
  49. package/.claude/agents/optimization/load-balancer.md +430 -430
  50. package/.claude/agents/optimization/performance-monitor.md +671 -671
  51. package/.claude/agents/optimization/resource-allocator.md +673 -673
  52. package/.claude/agents/optimization/topology-optimizer.md +807 -807
  53. package/.claude/agents/payments/agentic-payments.md +126 -126
  54. package/.claude/agents/sona/sona-learning-optimizer.md +74 -74
  55. package/.claude/agents/sparc/architecture.md +698 -698
  56. package/.claude/agents/sparc/pseudocode.md +519 -519
  57. package/.claude/agents/sparc/refinement.md +801 -801
  58. package/.claude/agents/sparc/specification.md +477 -477
  59. package/.claude/agents/specialized/mobile/spec-mobile-react-native.md +224 -224
  60. package/.claude/agents/specialized/spec-mobile-react-native.md +226 -226
  61. package/.claude/agents/sublinear/consensus-coordinator.md +337 -337
  62. package/.claude/agents/sublinear/matrix-optimizer.md +184 -184
  63. package/.claude/agents/sublinear/pagerank-analyzer.md +298 -298
  64. package/.claude/agents/sublinear/performance-optimizer.md +367 -367
  65. package/.claude/agents/sublinear/trading-predictor.md +245 -245
  66. package/.claude/agents/swarm/adaptive-coordinator.md +1126 -1126
  67. package/.claude/agents/swarm/hierarchical-coordinator.md +709 -709
  68. package/.claude/agents/swarm/mesh-coordinator.md +962 -962
  69. package/.claude/agents/templates/automation-smart-agent.md +204 -204
  70. package/.claude/agents/templates/base-template-generator.md +289 -289
  71. package/.claude/agents/templates/coordinator-swarm-init.md +89 -89
  72. package/.claude/agents/templates/github-pr-manager.md +176 -176
  73. package/.claude/agents/templates/implementer-sparc-coder.md +258 -258
  74. package/.claude/agents/templates/memory-coordinator.md +186 -186
  75. package/.claude/agents/templates/orchestrator-task.md +138 -138
  76. package/.claude/agents/templates/performance-analyzer.md +198 -198
  77. package/.claude/agents/templates/sparc-coordinator.md +513 -513
  78. package/.claude/agents/testing/production-validator.md +394 -394
  79. package/.claude/agents/testing/tdd-london-swarm.md +243 -243
  80. package/.claude/agents/v3/aidefence-guardian.md +282 -282
  81. package/.claude/agents/v3/claims-authorizer.md +208 -208
  82. package/.claude/agents/v3/collective-intelligence-coordinator.md +993 -993
  83. package/.claude/agents/v3/ddd-domain-expert.md +220 -220
  84. package/.claude/agents/v3/injection-analyst.md +236 -236
  85. package/.claude/agents/v3/performance-engineer.md +1233 -1233
  86. package/.claude/agents/v3/pii-detector.md +151 -151
  87. package/.claude/agents/v3/reasoningbank-learner.md +213 -213
  88. package/.claude/agents/v3/security-architect-aidefence.md +410 -410
  89. package/.claude/agents/v3/security-architect.md +867 -867
  90. package/.claude/agents/v3/swarm-memory-manager.md +157 -157
  91. package/.claude/agents/v3/v3-integration-architect.md +205 -205
  92. package/.claude/commands/agents/README.md +50 -50
  93. package/.claude/commands/agents/agent-capabilities.md +140 -140
  94. package/.claude/commands/agents/agent-coordination.md +28 -28
  95. package/.claude/commands/agents/agent-spawning.md +28 -28
  96. package/.claude/commands/agents/agent-types.md +216 -216
  97. package/.claude/commands/agents/health.md +139 -139
  98. package/.claude/commands/agents/list.md +100 -100
  99. package/.claude/commands/agents/logs.md +130 -130
  100. package/.claude/commands/agents/metrics.md +122 -122
  101. package/.claude/commands/agents/pool.md +127 -127
  102. package/.claude/commands/agents/spawn.md +140 -140
  103. package/.claude/commands/agents/status.md +115 -115
  104. package/.claude/commands/agents/stop.md +102 -102
  105. package/.claude/commands/analysis/COMMAND_COMPLIANCE_REPORT.md +53 -53
  106. package/.claude/commands/analysis/README.md +9 -9
  107. package/.claude/commands/analysis/bottleneck-detect.md +162 -162
  108. package/.claude/commands/analysis/performance-bottlenecks.md +58 -58
  109. package/.claude/commands/analysis/performance-report.md +25 -25
  110. package/.claude/commands/analysis/token-efficiency.md +44 -44
  111. package/.claude/commands/analysis/token-usage.md +25 -25
  112. package/.claude/commands/automation/README.md +9 -9
  113. package/.claude/commands/automation/auto-agent.md +122 -122
  114. package/.claude/commands/automation/self-healing.md +105 -105
  115. package/.claude/commands/automation/session-memory.md +89 -89
  116. package/.claude/commands/automation/smart-agents.md +72 -72
  117. package/.claude/commands/automation/smart-spawn.md +25 -25
  118. package/.claude/commands/automation/workflow-select.md +25 -25
  119. package/.claude/commands/claude-flow-help.md +103 -103
  120. package/.claude/commands/claude-flow-memory.md +107 -107
  121. package/.claude/commands/claude-flow-swarm.md +205 -205
  122. package/.claude/commands/coordination/README.md +9 -9
  123. package/.claude/commands/coordination/agent-spawn.md +25 -25
  124. package/.claude/commands/coordination/init.md +44 -44
  125. package/.claude/commands/coordination/orchestrate.md +43 -43
  126. package/.claude/commands/coordination/spawn.md +45 -45
  127. package/.claude/commands/coordination/swarm-init.md +85 -85
  128. package/.claude/commands/coordination/task-orchestrate.md +25 -25
  129. package/.claude/commands/github/README.md +11 -11
  130. package/.claude/commands/github/code-review-swarm.md +513 -513
  131. package/.claude/commands/github/code-review.md +25 -25
  132. package/.claude/commands/github/github-modes.md +146 -146
  133. package/.claude/commands/github/github-swarm.md +121 -121
  134. package/.claude/commands/github/issue-tracker.md +291 -291
  135. package/.claude/commands/github/issue-triage.md +25 -25
  136. package/.claude/commands/github/multi-repo-swarm.md +518 -518
  137. package/.claude/commands/github/pr-enhance.md +26 -26
  138. package/.claude/commands/github/pr-manager.md +169 -169
  139. package/.claude/commands/github/project-board-sync.md +470 -470
  140. package/.claude/commands/github/release-manager.md +339 -339
  141. package/.claude/commands/github/release-swarm.md +543 -543
  142. package/.claude/commands/github/repo-analyze.md +25 -25
  143. package/.claude/commands/github/repo-architect.md +366 -366
  144. package/.claude/commands/github/swarm-issue.md +484 -484
  145. package/.claude/commands/github/swarm-pr.md +287 -287
  146. package/.claude/commands/github/sync-coordinator.md +302 -302
  147. package/.claude/commands/github/workflow-automation.md +441 -441
  148. package/.claude/commands/hive-mind/README.md +17 -17
  149. package/.claude/commands/hive-mind/hive-mind-consensus.md +8 -8
  150. package/.claude/commands/hive-mind/hive-mind-init.md +18 -18
  151. package/.claude/commands/hive-mind/hive-mind-memory.md +8 -8
  152. package/.claude/commands/hive-mind/hive-mind-metrics.md +8 -8
  153. package/.claude/commands/hive-mind/hive-mind-resume.md +8 -8
  154. package/.claude/commands/hive-mind/hive-mind-sessions.md +8 -8
  155. package/.claude/commands/hive-mind/hive-mind-spawn.md +21 -21
  156. package/.claude/commands/hive-mind/hive-mind-status.md +8 -8
  157. package/.claude/commands/hive-mind/hive-mind-stop.md +8 -8
  158. package/.claude/commands/hive-mind/hive-mind-wizard.md +8 -8
  159. package/.claude/commands/hive-mind/hive-mind.md +27 -27
  160. package/.claude/commands/hooks/README.md +11 -11
  161. package/.claude/commands/hooks/overview.md +57 -57
  162. package/.claude/commands/hooks/post-edit.md +117 -117
  163. package/.claude/commands/hooks/post-task.md +112 -112
  164. package/.claude/commands/hooks/pre-edit.md +113 -113
  165. package/.claude/commands/hooks/pre-task.md +111 -111
  166. package/.claude/commands/hooks/session-end.md +118 -118
  167. package/.claude/commands/hooks/setup.md +102 -102
  168. package/.claude/commands/memory/README.md +9 -9
  169. package/.claude/commands/memory/memory-persist.md +25 -25
  170. package/.claude/commands/memory/memory-search.md +25 -25
  171. package/.claude/commands/memory/memory-usage.md +25 -25
  172. package/.claude/commands/memory/neural.md +47 -47
  173. package/.claude/commands/monitoring/README.md +9 -9
  174. package/.claude/commands/monitoring/agent-metrics.md +25 -25
  175. package/.claude/commands/monitoring/agents.md +44 -44
  176. package/.claude/commands/monitoring/real-time-view.md +25 -25
  177. package/.claude/commands/monitoring/status.md +46 -46
  178. package/.claude/commands/monitoring/swarm-monitor.md +25 -25
  179. package/.claude/commands/optimization/README.md +9 -9
  180. package/.claude/commands/optimization/auto-topology.md +61 -61
  181. package/.claude/commands/optimization/cache-manage.md +25 -25
  182. package/.claude/commands/optimization/parallel-execute.md +25 -25
  183. package/.claude/commands/optimization/parallel-execution.md +49 -49
  184. package/.claude/commands/optimization/topology-optimize.md +25 -25
  185. package/.claude/commands/pair/README.md +260 -260
  186. package/.claude/commands/pair/commands.md +545 -545
  187. package/.claude/commands/pair/config.md +509 -509
  188. package/.claude/commands/pair/examples.md +511 -511
  189. package/.claude/commands/pair/modes.md +347 -347
  190. package/.claude/commands/pair/session.md +406 -406
  191. package/.claude/commands/pair/start.md +208 -208
  192. package/.claude/commands/sparc/analyzer.md +51 -51
  193. package/.claude/commands/sparc/architect.md +53 -53
  194. package/.claude/commands/sparc/ask.md +97 -97
  195. package/.claude/commands/sparc/batch-executor.md +54 -54
  196. package/.claude/commands/sparc/code.md +89 -89
  197. package/.claude/commands/sparc/coder.md +54 -54
  198. package/.claude/commands/sparc/debug.md +83 -83
  199. package/.claude/commands/sparc/debugger.md +54 -54
  200. package/.claude/commands/sparc/designer.md +53 -53
  201. package/.claude/commands/sparc/devops.md +109 -109
  202. package/.claude/commands/sparc/docs-writer.md +80 -80
  203. package/.claude/commands/sparc/documenter.md +54 -54
  204. package/.claude/commands/sparc/innovator.md +54 -54
  205. package/.claude/commands/sparc/integration.md +83 -83
  206. package/.claude/commands/sparc/mcp.md +117 -117
  207. package/.claude/commands/sparc/memory-manager.md +54 -54
  208. package/.claude/commands/sparc/optimizer.md +54 -54
  209. package/.claude/commands/sparc/orchestrator.md +131 -131
  210. package/.claude/commands/sparc/post-deployment-monitoring-mode.md +83 -83
  211. package/.claude/commands/sparc/refinement-optimization-mode.md +83 -83
  212. package/.claude/commands/sparc/researcher.md +54 -54
  213. package/.claude/commands/sparc/reviewer.md +54 -54
  214. package/.claude/commands/sparc/security-review.md +80 -80
  215. package/.claude/commands/sparc/sparc-modes.md +174 -174
  216. package/.claude/commands/sparc/sparc.md +111 -111
  217. package/.claude/commands/sparc/spec-pseudocode.md +80 -80
  218. package/.claude/commands/sparc/supabase-admin.md +348 -348
  219. package/.claude/commands/sparc/swarm-coordinator.md +54 -54
  220. package/.claude/commands/sparc/tdd.md +54 -54
  221. package/.claude/commands/sparc/tester.md +54 -54
  222. package/.claude/commands/sparc/tutorial.md +79 -79
  223. package/.claude/commands/sparc/workflow-manager.md +54 -54
  224. package/.claude/commands/sparc.md +166 -166
  225. package/.claude/commands/stream-chain/pipeline.md +120 -120
  226. package/.claude/commands/stream-chain/run.md +69 -69
  227. package/.claude/commands/swarm/README.md +15 -15
  228. package/.claude/commands/swarm/analysis.md +95 -95
  229. package/.claude/commands/swarm/development.md +96 -96
  230. package/.claude/commands/swarm/examples.md +168 -168
  231. package/.claude/commands/swarm/maintenance.md +102 -102
  232. package/.claude/commands/swarm/optimization.md +117 -117
  233. package/.claude/commands/swarm/research.md +136 -136
  234. package/.claude/commands/swarm/swarm-analysis.md +8 -8
  235. package/.claude/commands/swarm/swarm-background.md +8 -8
  236. package/.claude/commands/swarm/swarm-init.md +19 -19
  237. package/.claude/commands/swarm/swarm-modes.md +8 -8
  238. package/.claude/commands/swarm/swarm-monitor.md +8 -8
  239. package/.claude/commands/swarm/swarm-spawn.md +19 -19
  240. package/.claude/commands/swarm/swarm-status.md +8 -8
  241. package/.claude/commands/swarm/swarm-strategies.md +8 -8
  242. package/.claude/commands/swarm/swarm.md +87 -87
  243. package/.claude/commands/swarm/testing.md +131 -131
  244. package/.claude/commands/training/README.md +9 -9
  245. package/.claude/commands/training/model-update.md +25 -25
  246. package/.claude/commands/training/neural-patterns.md +107 -107
  247. package/.claude/commands/training/neural-train.md +75 -75
  248. package/.claude/commands/training/pattern-learn.md +25 -25
  249. package/.claude/commands/training/specialization.md +62 -62
  250. package/.claude/commands/truth/start.md +142 -142
  251. package/.claude/commands/verify/check.md +49 -49
  252. package/.claude/commands/verify/start.md +127 -127
  253. package/.claude/commands/workflows/README.md +9 -9
  254. package/.claude/commands/workflows/development.md +77 -77
  255. package/.claude/commands/workflows/research.md +62 -62
  256. package/.claude/commands/workflows/workflow-create.md +25 -25
  257. package/.claude/commands/workflows/workflow-execute.md +25 -25
  258. package/.claude/commands/workflows/workflow-export.md +25 -25
  259. package/.claude/eval/human-relevance-frozen-v1.json +17 -17
  260. package/.claude/evolve-proof/generation-0.json +211 -211
  261. package/.claude/evolve-proof/real-generation-0.json +406 -406
  262. package/.claude/evolve-proof/real-generation-1.json +406 -406
  263. package/.claude/helpers/.helpers-version +1 -1
  264. package/.claude/helpers/README.md +96 -96
  265. package/.claude/helpers/adr-compliance.sh +186 -186
  266. package/.claude/helpers/auto-commit.sh +178 -178
  267. package/.claude/helpers/auto-memory-hook.mjs +0 -0
  268. package/.claude/helpers/checkpoint-manager.sh +251 -251
  269. package/.claude/helpers/daemon-manager.sh +252 -252
  270. package/.claude/helpers/ddd-tracker.sh +144 -144
  271. package/.claude/helpers/github-safe.js +156 -156
  272. package/.claude/helpers/github-setup.sh +45 -45
  273. package/.claude/helpers/guidance-hook.sh +13 -13
  274. package/.claude/helpers/guidance-hooks.sh +102 -102
  275. package/.claude/helpers/health-monitor.sh +108 -108
  276. package/.claude/helpers/helpers.manifest.json +2 -2
  277. package/.claude/helpers/hook-handler.cjs +0 -0
  278. package/.claude/helpers/intelligence.cjs +0 -0
  279. package/.claude/helpers/learning-hooks.sh +329 -329
  280. package/.claude/helpers/learning-optimizer.sh +127 -127
  281. package/.claude/helpers/learning-service.mjs +1144 -1144
  282. package/.claude/helpers/memory.js +83 -83
  283. package/.claude/helpers/metrics-db.mjs +503 -503
  284. package/.claude/helpers/pattern-consolidator.sh +86 -86
  285. package/.claude/helpers/perf-worker.sh +160 -160
  286. package/.claude/helpers/post-commit +16 -16
  287. package/.claude/helpers/pre-commit +26 -26
  288. package/.claude/helpers/quick-start.sh +19 -19
  289. package/.claude/helpers/router.js +105 -105
  290. package/.claude/helpers/security-scanner.sh +127 -127
  291. package/.claude/helpers/session.js +157 -157
  292. package/.claude/helpers/setup-mcp.sh +18 -18
  293. package/.claude/helpers/standard-checkpoint-hooks.sh +189 -189
  294. package/.claude/helpers/statusline-hook.sh +21 -21
  295. package/.claude/helpers/statusline.cjs +0 -0
  296. package/.claude/helpers/statusline.js +340 -340
  297. package/.claude/helpers/swarm-comms.sh +353 -353
  298. package/.claude/helpers/swarm-hooks.sh +761 -761
  299. package/.claude/helpers/swarm-monitor.sh +210 -210
  300. package/.claude/helpers/sync-v3-metrics.sh +245 -245
  301. package/.claude/helpers/update-v3-progress.sh +165 -165
  302. package/.claude/helpers/v3-quick-status.sh +57 -57
  303. package/.claude/helpers/v3.sh +110 -110
  304. package/.claude/helpers/validate-v3-config.sh +215 -215
  305. package/.claude/helpers/worker-manager.sh +170 -170
  306. package/.claude/proven-config.json +42 -0
  307. package/.claude/proven-config.manifest.json +37 -37
  308. package/.claude/proven-config.signed.json +41 -41
  309. package/.claude/settings.json +182 -182
  310. package/.claude/skills/agentdb-advanced/SKILL.md +550 -550
  311. package/.claude/skills/agentdb-learning/SKILL.md +545 -545
  312. package/.claude/skills/agentdb-memory-patterns/SKILL.md +339 -339
  313. package/.claude/skills/agentdb-optimization/SKILL.md +509 -509
  314. package/.claude/skills/agentdb-vector-search/SKILL.md +339 -339
  315. package/.claude/skills/browser/SKILL.md +204 -204
  316. package/.claude/skills/dual-mode/README.md +71 -71
  317. package/.claude/skills/dual-mode/dual-collect.md +103 -103
  318. package/.claude/skills/dual-mode/dual-coordinate.md +85 -85
  319. package/.claude/skills/dual-mode/dual-spawn.md +81 -81
  320. package/.claude/skills/flow-nexus-neural/SKILL.md +727 -727
  321. package/.claude/skills/flow-nexus-platform/SKILL.md +1154 -1154
  322. package/.claude/skills/flow-nexus-swarm/SKILL.md +604 -604
  323. package/.claude/skills/github-code-review/SKILL.md +1125 -1125
  324. package/.claude/skills/github-multi-repo/SKILL.md +862 -862
  325. package/.claude/skills/github-project-management/SKILL.md +1262 -1262
  326. package/.claude/skills/github-release-management/SKILL.md +1064 -1064
  327. package/.claude/skills/github-workflow-automation/SKILL.md +1047 -1047
  328. package/.claude/skills/hooks-automation/SKILL.md +1201 -1201
  329. package/.claude/skills/pair-programming/SKILL.md +1202 -1202
  330. package/.claude/skills/reasoningbank-agentdb/SKILL.md +446 -446
  331. package/.claude/skills/reasoningbank-intelligence/SKILL.md +201 -201
  332. package/.claude/skills/skill-builder/SKILL.md +910 -910
  333. package/.claude/skills/sparc-methodology/SKILL.md +1106 -1106
  334. package/.claude/skills/stream-chain/SKILL.md +560 -560
  335. package/.claude/skills/swarm-advanced/SKILL.md +970 -970
  336. package/.claude/skills/swarm-orchestration/SKILL.md +179 -179
  337. package/.claude/skills/v3-cli-modernization/SKILL.md +871 -871
  338. package/.claude/skills/v3-core-implementation/SKILL.md +796 -796
  339. package/.claude/skills/v3-ddd-architecture/SKILL.md +441 -441
  340. package/.claude/skills/v3-integration-deep/SKILL.md +240 -240
  341. package/.claude/skills/v3-mcp-optimization/SKILL.md +776 -776
  342. package/.claude/skills/v3-memory-unification/SKILL.md +173 -173
  343. package/.claude/skills/v3-performance-optimization/SKILL.md +389 -389
  344. package/.claude/skills/v3-security-overhaul/SKILL.md +81 -81
  345. package/.claude/skills/v3-swarm-coordination/SKILL.md +339 -339
  346. package/.claude/skills/verification-quality/SKILL.md +691 -691
  347. package/README.md +419 -419
  348. package/bin/cli.js +314 -314
  349. package/bin/mcp-server.js +224 -224
  350. package/bin/preinstall.cjs +2 -2
  351. package/catalog-manifest.json +2 -2
  352. package/dist/src/autopilot-state.js +24 -7
  353. package/dist/src/benchmarks/gaia-critic.js +24 -24
  354. package/dist/src/business-pods/bbs-budget-tracker.js +53 -53
  355. package/dist/src/commands/completions.js +409 -409
  356. package/dist/src/commands/daemon.js +44 -44
  357. package/dist/src/commands/embeddings.js +26 -26
  358. package/dist/src/commands/hive-mind.js +97 -97
  359. package/dist/src/commands/hooks.js +31 -10
  360. package/dist/src/commands/init.js +202 -34
  361. package/dist/src/commands/memory.js +12 -1
  362. package/dist/src/commands/ruvector/backup.js +23 -23
  363. package/dist/src/commands/ruvector/benchmark.js +31 -31
  364. package/dist/src/commands/ruvector/import.js +14 -14
  365. package/dist/src/commands/ruvector/init.js +115 -115
  366. package/dist/src/commands/ruvector/migrate.js +99 -99
  367. package/dist/src/commands/ruvector/optimize.js +51 -51
  368. package/dist/src/commands/ruvector/setup.js +624 -624
  369. package/dist/src/commands/ruvector/status.js +38 -38
  370. package/dist/src/config/proven-config.js +2 -2
  371. package/dist/src/funnel/disclosure.js +13 -2
  372. package/dist/src/funnel/messages.d.ts +12 -10
  373. package/dist/src/funnel/messages.js +83 -11
  374. package/dist/src/init/claudemd-generator.js +231 -231
  375. package/dist/src/init/executor.js +453 -453
  376. package/dist/src/init/helper-signing.js +2 -2
  377. package/dist/src/init/helpers-generator.js +751 -751
  378. package/dist/src/init/statusline-generator.js +24 -24
  379. package/dist/src/mcp-tools/agentdb-tools.js +15 -15
  380. package/dist/src/mcp-tools/browser-intent-tools.js +19 -19
  381. package/dist/src/mcp-tools/browser-tools.js +8 -0
  382. package/dist/src/mcp-tools/hooks-tools.js +21 -0
  383. package/dist/src/mcp-tools/memory-tools.js +4 -3
  384. package/dist/src/memory/graph-edge-writer.js +22 -22
  385. package/dist/src/memory/memory-bridge.js +248 -158
  386. package/dist/src/memory/memory-initializer.js +407 -407
  387. package/dist/src/memory/rabitq-index.js +5 -5
  388. package/dist/src/parser.js +25 -9
  389. package/dist/src/proxy/verify.js +2 -2
  390. package/dist/src/runtime/headless.js +28 -28
  391. package/dist/src/services/distill-tuning.js +7 -7
  392. package/dist/src/services/headless-worker-executor.js +84 -84
  393. package/dist/src/services/memory-distillation.js +4 -4
  394. package/dist/src/services/worker-daemon.js +7 -4
  395. package/dist/src/transfer/deploy-seraphine.js +23 -23
  396. package/package.json +137 -137
  397. package/plugins/ruflo-metaharness/.claude-plugin/plugin.json +32 -32
  398. package/plugins/ruflo-metaharness/README.md +72 -72
  399. package/plugins/ruflo-metaharness/agents/metaharness-architect.md +58 -58
  400. package/plugins/ruflo-metaharness/commands/ruflo-metaharness.md +48 -48
  401. package/plugins/ruflo-metaharness/scripts/_darwin.mjs +210 -210
  402. package/plugins/ruflo-metaharness/scripts/_harness.mjs +330 -330
  403. package/plugins/ruflo-metaharness/scripts/_invoke.mjs +231 -231
  404. package/plugins/ruflo-metaharness/scripts/_redblue.mjs +143 -143
  405. package/plugins/ruflo-metaharness/scripts/_similarity.mjs +161 -161
  406. package/plugins/ruflo-metaharness/scripts/_spike-similarity.mjs +223 -223
  407. package/plugins/ruflo-metaharness/scripts/audit-list.mjs +158 -158
  408. package/plugins/ruflo-metaharness/scripts/audit-trend.mjs +272 -272
  409. package/plugins/ruflo-metaharness/scripts/bench-parse-mcp-scan.mjs +146 -146
  410. package/plugins/ruflo-metaharness/scripts/bench-recordpair-overhead.mjs +186 -186
  411. package/plugins/ruflo-metaharness/scripts/bench-similarity.mjs +177 -177
  412. package/plugins/ruflo-metaharness/scripts/bench.mjs +95 -95
  413. package/plugins/ruflo-metaharness/scripts/drift-from-history.mjs +363 -363
  414. package/plugins/ruflo-metaharness/scripts/evolve.mjs +404 -404
  415. package/plugins/ruflo-metaharness/scripts/genome.mjs +80 -80
  416. package/plugins/ruflo-metaharness/scripts/gepa.mjs +153 -153
  417. package/plugins/ruflo-metaharness/scripts/learn.mjs +127 -127
  418. package/plugins/ruflo-metaharness/scripts/mcp-scan.mjs +111 -111
  419. package/plugins/ruflo-metaharness/scripts/mint.mjs +126 -126
  420. package/plugins/ruflo-metaharness/scripts/oia-audit.mjs +228 -228
  421. package/plugins/ruflo-metaharness/scripts/redblue.mjs +286 -286
  422. package/plugins/ruflo-metaharness/scripts/router-parallel-analyze.mjs +250 -250
  423. package/plugins/ruflo-metaharness/scripts/score.mjs +92 -92
  424. package/plugins/ruflo-metaharness/scripts/security-bench.mjs +174 -174
  425. package/plugins/ruflo-metaharness/scripts/similarity.mjs +158 -158
  426. package/plugins/ruflo-metaharness/scripts/smoke.sh +2356 -2356
  427. package/plugins/ruflo-metaharness/scripts/test-graceful-degradation.mjs +165 -165
  428. package/plugins/ruflo-metaharness/scripts/test-mcp-tools.mjs +472 -472
  429. package/plugins/ruflo-metaharness/scripts/test-parallel-pipeline.mjs +204 -204
  430. package/plugins/ruflo-metaharness/scripts/test-pipeline-roundtrip.mjs +586 -586
  431. package/plugins/ruflo-metaharness/scripts/test-similarity.mjs +334 -334
  432. package/plugins/ruflo-metaharness/scripts/test-with-openrouter.mjs +229 -229
  433. package/plugins/ruflo-metaharness/scripts/threat-model.mjs +59 -59
  434. package/plugins/ruflo-metaharness/skills/harness-bench/SKILL.md +64 -64
  435. package/plugins/ruflo-metaharness/skills/harness-drift-from-history/SKILL.md +65 -65
  436. package/plugins/ruflo-metaharness/skills/harness-evolve/SKILL.md +131 -131
  437. package/plugins/ruflo-metaharness/skills/harness-genome/SKILL.md +54 -54
  438. package/plugins/ruflo-metaharness/skills/harness-gepa/SKILL.md +65 -65
  439. package/plugins/ruflo-metaharness/skills/harness-learn/SKILL.md +65 -65
  440. package/plugins/ruflo-metaharness/skills/harness-mcp-scan/SKILL.md +49 -49
  441. package/plugins/ruflo-metaharness/skills/harness-mint/SKILL.md +72 -72
  442. package/plugins/ruflo-metaharness/skills/harness-oia-audit/SKILL.md +79 -79
  443. package/plugins/ruflo-metaharness/skills/harness-score/SKILL.md +66 -66
  444. package/plugins/ruflo-metaharness/skills/harness-security-bench/SKILL.md +101 -101
  445. package/plugins/ruflo-metaharness/skills/harness-similarity/SKILL.md +67 -67
  446. package/plugins/ruflo-metaharness/skills/harness-threat-model/SKILL.md +41 -41
  447. package/scripts/postinstall.cjs +153 -153
@@ -1,64 +1,64 @@
1
- ---
2
- name: harness-bench
3
- description: Manage `@metaharness/darwin` bench suites — `bench create <repo>` scaffolds a JSON suite from a repo's test corpus; `bench verify <suite.json>` checks suite well-formedness. Bench suites are the fixed evaluation corpora that `harness-evolve --bench <suite.json>` scores variants against, decoupling evolution from the repo's natural tests. Degrades gracefully when @metaharness/darwin is absent.
4
- argument-hint: "--op create --repo <path> [--out <path>] | --op verify --suite <path>"
5
- allowed-tools: Bash
6
- ---
7
-
8
- Surfaces `metaharness-darwin bench <create|verify>` — the supporting verb
9
- for `harness-evolve --bench`. Use when you want evolution scored against a
10
- fixed corpus (independent of `npm test`) so champion fitness is comparable
11
- across commits or across forks of the same harness.
12
-
13
- ## When to use
14
-
15
- - Setting up a new evolution pipeline for a repo whose `npm test` is
16
- flaky, slow, or undersized — scaffold a deterministic bench suite once,
17
- then evolve against it repeatedly.
18
- - CI: `bench verify` the checked-in suite on every PR that touches it
19
- (cheap; ~5s).
20
- - Forking a harness to a new domain: copy and edit the suite to retarget
21
- the evaluation without losing comparability to the parent.
22
-
23
- ## Algorithm
24
-
25
- Implementation: [`scripts/bench.mjs`](../../scripts/bench.mjs).
26
-
27
- ### `--op create`
28
- 1. Resolve `--repo` path; reject if missing.
29
- 2. Shell to `metaharness-darwin bench create <repo> [--out <suite.json>]`.
30
- 3. Default output path: `<repo>/.metaharness/bench/suite.json` (chosen by upstream).
31
- 4. Suite shape (per upstream): array of `{ input, expectedOutput, weight }` tasks
32
- derived from existing test cases.
33
-
34
- ### `--op verify`
35
- 1. Resolve `--suite` path; reject if missing.
36
- 2. Shell to `metaharness-darwin bench verify <suite.json>`.
37
- 3. Exit 1 if any task malformed (upstream's signal).
38
-
39
- ## Output shape
40
-
41
- ```json
42
- {
43
- "success": true,
44
- "data": {
45
- "op": "verify",
46
- "taskCount": 42,
47
- "wellFormed": true,
48
- "durationMs": 870
49
- }
50
- }
51
- ```
52
-
53
- ## Exit codes
54
-
55
- | Code | Meaning |
56
- |---|---|
57
- | 0 | OK (or degraded — Darwin absent) |
58
- | 1 | `--op verify` and suite malformed |
59
- | 2 | Config error or upstream invocation failure |
60
-
61
- ## Graceful degradation
62
-
63
- When `@metaharness/darwin` is absent, emits the standard `{degraded: true,
64
- reason: 'metaharness-darwin-not-available'}` payload and exits 0.
1
+ ---
2
+ name: harness-bench
3
+ description: Manage `@metaharness/darwin` bench suites — `bench create <repo>` scaffolds a JSON suite from a repo's test corpus; `bench verify <suite.json>` checks suite well-formedness. Bench suites are the fixed evaluation corpora that `harness-evolve --bench <suite.json>` scores variants against, decoupling evolution from the repo's natural tests. Degrades gracefully when @metaharness/darwin is absent.
4
+ argument-hint: "--op create --repo <path> [--out <path>] | --op verify --suite <path>"
5
+ allowed-tools: Bash
6
+ ---
7
+
8
+ Surfaces `metaharness-darwin bench <create|verify>` — the supporting verb
9
+ for `harness-evolve --bench`. Use when you want evolution scored against a
10
+ fixed corpus (independent of `npm test`) so champion fitness is comparable
11
+ across commits or across forks of the same harness.
12
+
13
+ ## When to use
14
+
15
+ - Setting up a new evolution pipeline for a repo whose `npm test` is
16
+ flaky, slow, or undersized — scaffold a deterministic bench suite once,
17
+ then evolve against it repeatedly.
18
+ - CI: `bench verify` the checked-in suite on every PR that touches it
19
+ (cheap; ~5s).
20
+ - Forking a harness to a new domain: copy and edit the suite to retarget
21
+ the evaluation without losing comparability to the parent.
22
+
23
+ ## Algorithm
24
+
25
+ Implementation: [`scripts/bench.mjs`](../../scripts/bench.mjs).
26
+
27
+ ### `--op create`
28
+ 1. Resolve `--repo` path; reject if missing.
29
+ 2. Shell to `metaharness-darwin bench create <repo> [--out <suite.json>]`.
30
+ 3. Default output path: `<repo>/.metaharness/bench/suite.json` (chosen by upstream).
31
+ 4. Suite shape (per upstream): array of `{ input, expectedOutput, weight }` tasks
32
+ derived from existing test cases.
33
+
34
+ ### `--op verify`
35
+ 1. Resolve `--suite` path; reject if missing.
36
+ 2. Shell to `metaharness-darwin bench verify <suite.json>`.
37
+ 3. Exit 1 if any task malformed (upstream's signal).
38
+
39
+ ## Output shape
40
+
41
+ ```json
42
+ {
43
+ "success": true,
44
+ "data": {
45
+ "op": "verify",
46
+ "taskCount": 42,
47
+ "wellFormed": true,
48
+ "durationMs": 870
49
+ }
50
+ }
51
+ ```
52
+
53
+ ## Exit codes
54
+
55
+ | Code | Meaning |
56
+ |---|---|
57
+ | 0 | OK (or degraded — Darwin absent) |
58
+ | 1 | `--op verify` and suite malformed |
59
+ | 2 | Config error or upstream invocation failure |
60
+
61
+ ## Graceful degradation
62
+
63
+ When `@metaharness/darwin` is absent, emits the standard `{degraded: true,
64
+ reason: 'metaharness-darwin-not-available'}` payload and exits 0.
@@ -1,65 +1,65 @@
1
- ---
2
- name: harness-drift-from-history
3
- description: One-command drift detection. Composes audit-list + oia-audit + audit-trend into a single primitive — finds the most recent audit in `metaharness-audit` namespace, runs a fresh audit against the current repo, diffs them via ADR-152 §3.1 similarity, and alerts when structural distance crosses `--threshold`. Iter 53 of ADR-150 deep integration.
4
- argument-hint: "[--path .] [--baseline-since 7d] [--threshold 0.95] [--dry-run] [--format json|table]"
5
- allowed-tools: Bash
6
- ---
7
-
8
- The natural ops question after running `oia-audit` weekly is "did anything drift?" Before this skill, the answer required a three-step sequence:
9
-
10
- ```bash
11
- npx ruflo metaharness audit-list --format json # → pick a key by hand
12
- npx ruflo metaharness oia-audit --format json > /tmp/curr.json
13
- npx ruflo metaharness audit-trend \
14
- --baseline-key <picked-key> --current /tmp/curr.json \
15
- --alert-on-distance-below 0.95
16
- ```
17
-
18
- This skill collapses it into one command:
19
-
20
- ```bash
21
- npx ruflo metaharness drift-from-history --threshold 0.95
22
- ```
23
-
24
- ## What it does
25
-
26
- 1. Lists records from `metaharness-audit` namespace via audit-list.mjs
27
- 2. Picks the most recent record by `startedAt` (or `--baseline-since 7d` skips anything newer than 7 days)
28
- 3. Runs a fresh `oia-audit` against the current path
29
- 4. Diffs the two via audit-trend, applying `--alert-on-distance-below ${threshold}`
30
- 5. Returns the structured drift report
31
-
32
- ## Architectural constraint inheritance (ADR-150)
33
-
34
- | Constraint | How drift-from-history satisfies it |
35
- |---|---|
36
- | Removable | Pure subprocess composition over existing scripts — no new `@metaharness/*` import |
37
- | Optional | If oia-audit reports `degraded:true`, this skill exits 3 with a degraded payload |
38
- | Graceful | Empty audit history → exit 2 with hint to seed it; never crashes |
39
- | CI-gate | Smoke step 17z16 anchors the dispatcher entry + subcommand listing |
40
-
41
- ## Exit codes
42
-
43
- - 0 — similarity ≥ threshold (or threshold not crossed)
44
- - 1 — drift detected: similarity < threshold (alert fired)
45
- - 2 — config error (no history, audit-list failed)
46
- - 3 — upstream metaharness absent (degraded payload returned)
47
-
48
- ## Example
49
-
50
- ```bash
51
- $ npx ruflo metaharness drift-from-history --threshold 0.95
52
- # drift-from-history
53
-
54
- Baseline: audit-2026-06-16T22-58-47-840Z
55
- Current: 2026-06-16T23:05:02.231Z
56
-
57
- Structural similarity: 1 (near-identical)
58
- Distance: 0
59
-
60
- ✓ similarity ≥ 0.95 — OK
61
- ```
62
-
63
- ## Implementation
64
-
65
- [`scripts/drift-from-history.mjs`](../../scripts/drift-from-history.mjs)
1
+ ---
2
+ name: harness-drift-from-history
3
+ description: One-command drift detection. Composes audit-list + oia-audit + audit-trend into a single primitive — finds the most recent audit in `metaharness-audit` namespace, runs a fresh audit against the current repo, diffs them via ADR-152 §3.1 similarity, and alerts when structural distance crosses `--threshold`. Iter 53 of ADR-150 deep integration.
4
+ argument-hint: "[--path .] [--baseline-since 7d] [--threshold 0.95] [--dry-run] [--format json|table]"
5
+ allowed-tools: Bash
6
+ ---
7
+
8
+ The natural ops question after running `oia-audit` weekly is "did anything drift?" Before this skill, the answer required a three-step sequence:
9
+
10
+ ```bash
11
+ npx ruflo metaharness audit-list --format json # → pick a key by hand
12
+ npx ruflo metaharness oia-audit --format json > /tmp/curr.json
13
+ npx ruflo metaharness audit-trend \
14
+ --baseline-key <picked-key> --current /tmp/curr.json \
15
+ --alert-on-distance-below 0.95
16
+ ```
17
+
18
+ This skill collapses it into one command:
19
+
20
+ ```bash
21
+ npx ruflo metaharness drift-from-history --threshold 0.95
22
+ ```
23
+
24
+ ## What it does
25
+
26
+ 1. Lists records from `metaharness-audit` namespace via audit-list.mjs
27
+ 2. Picks the most recent record by `startedAt` (or `--baseline-since 7d` skips anything newer than 7 days)
28
+ 3. Runs a fresh `oia-audit` against the current path
29
+ 4. Diffs the two via audit-trend, applying `--alert-on-distance-below ${threshold}`
30
+ 5. Returns the structured drift report
31
+
32
+ ## Architectural constraint inheritance (ADR-150)
33
+
34
+ | Constraint | How drift-from-history satisfies it |
35
+ |---|---|
36
+ | Removable | Pure subprocess composition over existing scripts — no new `@metaharness/*` import |
37
+ | Optional | If oia-audit reports `degraded:true`, this skill exits 3 with a degraded payload |
38
+ | Graceful | Empty audit history → exit 2 with hint to seed it; never crashes |
39
+ | CI-gate | Smoke step 17z16 anchors the dispatcher entry + subcommand listing |
40
+
41
+ ## Exit codes
42
+
43
+ - 0 — similarity ≥ threshold (or threshold not crossed)
44
+ - 1 — drift detected: similarity < threshold (alert fired)
45
+ - 2 — config error (no history, audit-list failed)
46
+ - 3 — upstream metaharness absent (degraded payload returned)
47
+
48
+ ## Example
49
+
50
+ ```bash
51
+ $ npx ruflo metaharness drift-from-history --threshold 0.95
52
+ # drift-from-history
53
+
54
+ Baseline: audit-2026-06-16T22-58-47-840Z
55
+ Current: 2026-06-16T23:05:02.231Z
56
+
57
+ Structural similarity: 1 (near-identical)
58
+ Distance: 0
59
+
60
+ ✓ similarity ≥ 0.95 — OK
61
+ ```
62
+
63
+ ## Implementation
64
+
65
+ [`scripts/drift-from-history.mjs`](../../scripts/drift-from-history.mjs)
@@ -1,131 +1,131 @@
1
- ---
2
- name: harness-evolve
3
- description: Run `@metaharness/darwin evolve <repo>` to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and promote only measured wins. The model is frozen; the harness evolves. Closes the loop ADR-150 opens (score+genome describe; evolve changes). Degrades gracefully when @metaharness/darwin is absent (ADR-150 + ADR-153 architectural constraints).
4
- argument-hint: "--repo <path> [--generations 3] [--children 3] [--concurrency 2] [--sandbox real|mock|agent] [--selection pareto|quality-diversity|...] [--mutator deterministic|ruvllm] [--diagnose] [--confirm]"
5
- allowed-tools: Bash
6
- ---
7
-
8
- Surfaces the upstream `metaharness-darwin evolve` CLI as a ruflo skill. The
9
- **write** layer that pairs with ADR-150's read layer (score / genome /
10
- mcp-scan / threat-model / oia-audit). Use when you have a harness whose
11
- readiness scores are flat and you want to discover *which* surface mutation
12
- moves them — without retraining the foundation model.
13
-
14
- ## When to use
15
-
16
- - A `harness-score` result is below target and you don't know which policy
17
- surface is responsible.
18
- - You're seeding a harness for a new vertical and want to find a good
19
- starting configuration empirically rather than hand-tuning.
20
- - You're comparing your hand-tuned harness against an evolved baseline
21
- (treat darwin's champion as the strawman).
22
-
23
- ## When NOT to use
24
-
25
- - For continuous background optimization. Darwin Mode is human-initiated.
26
- Wire it into CI for one-shot exploration, not for autonomous self-modification.
27
- - For ruflo itself in CI. ADR-153 §5 explicitly rejects auto-evolving ruflo
28
- — the CI gate verifies graceful degradation, not convergence.
29
-
30
- ## Algorithm
31
-
32
- Implementation: [`scripts/evolve.mjs`](../../scripts/evolve.mjs).
33
-
34
- 1. Validate args (`--repo` exists, caps on `--generations` ≤ 50, `--children`
35
- ≤ 20, `--concurrency` ≤ 8, sandbox/selection/mutator are known values).
36
- 2. Without `--confirm`: print plan + exit 0 (mirrors `harness-mint` safety
37
- convention; defense in depth over the upstream `safety.ts` checks).
38
- 3. With `--confirm`: shell to `npx -y @metaharness/darwin@~0.8.0 metaharness-darwin evolve <repo> ...`
39
- via the shared `_darwin.mjs` async helper. Per-generation progress is
40
- forwarded to stderr; final champion JSON is captured from stdout.
41
- 4. Compute timeout from `generations × children × per-variant` (per-variant
42
- ≈ 60s real, ≈ 2s mock). Caller may override with `--timeout-ms`.
43
- 5. Honor upstream exit code 99 — propagate as "safety-disqualified", do not
44
- remap. This is a designed-in tripwire (a variant tripped `inspectVariant`
45
- for secrets / shell-out / network / dynamic-eval). See ADR-153 §"Safety model".
46
- 6. Optional `--alert-on-no-improvement`: exit 1 when champion ≤ parent.
47
-
48
- ## The seven mutation surfaces
49
-
50
- | Surface | What it owns |
51
- |---|---|
52
- | `planner` | task decomposition / step ordering |
53
- | `contextBuilder` | what gets fed into the prompt |
54
- | `reviewer` | self-critique / output verification |
55
- | `retryPolicy` | when + how to retry on failure |
56
- | `toolPolicy` | which tools the agent may use, under which conditions |
57
- | `memoryPolicy` | what to persist, recall, forget |
58
- | `scorePolicy` | how the agent grades its own output |
59
-
60
- One mutation per variant. Multi-surface mutations are not allowed (causal
61
- attribution stays clean).
62
-
63
- ## Output
64
-
65
- Reports land under `<repo>/.metaharness/`:
66
-
67
- ```
68
- .metaharness/
69
- archive.json # full lineage tree (sampling next gen draws from this)
70
- lineage.json # parent→child edges only
71
- variants/<id>/ # per-variant code (kept for audit)
72
- runs/<id>/ # per-variant sandbox test output
73
- reports/winner.json # final champion + score delta vs parent
74
- ```
75
-
76
- Skill stdout = JSON `{success, data: {champion, plan, durationMs, improved}}`
77
- (plus `data.diagnosis` when `--diagnose` is passed — see below).
78
-
79
- ## Failure diagnosis (`--diagnose`)
80
-
81
- GEPA's key trick is natural-language failure diagnosis from execution traces
82
- feeding the next mutation — not just scalar fitness. `--diagnose` adds a
83
- modest slice of that: after the evolution completes, the losing / failed
84
- variants' transcripts are run through darwin's GEPA library ops
85
- (`analyzeTranscript` + `classifyFailure`, via the shared `importGepa`
86
- resolver in `scripts/_darwin.mjs`) and a `diagnosis` section is appended to
87
- the emitted JSON:
88
-
89
- ```json
90
- "diagnosis": {
91
- "available": true,
92
- "scope": "losing-variants",
93
- "variants": [
94
- { "id": "g1_v0", "transcripts": 2,
95
- "failureClasses": { "exploration-loop": 1, "edit-mechanics": 1 },
96
- "dominantClass": "exploration-loop" }
97
- ],
98
- "totals": { "exploration-loop": 1, "edit-mechanics": 1 }
99
- }
100
- ```
101
-
102
- Upstream shape caveats (verified against `@metaharness/darwin@0.8.0`):
103
-
104
- - `metaharness-darwin evolve --json` prints a TEXT leaderboard — the stdout
105
- carries no JSON and no transcripts. Per-variant run records live at
106
- `<repo>/.metaharness/runs/<id>.json`.
107
- - Those run records hold sandbox exec traces (`{taskId, exitCode, stdout,
108
- stderr}`), which are NOT GEPA `{actionRaw, obs}` transcripts. Diagnosis
109
- therefore uses GEPA-shaped transcripts when a run record embeds them
110
- (agent sandbox / future upstream), falls back to the champion's transcript,
111
- and otherwise emits `diagnosis: {available: false, reason, traceSummary}`
112
- where `traceSummary` is a mechanical per-variant tally (tasks / failed /
113
- timedOut / blockedActions).
114
- - `--diagnose` NEVER fails the run — any internal error degrades to
115
- `{available: false, reason: "diagnosis-failed: ..."}`.
116
-
117
- ## Exit codes
118
-
119
- | Code | Meaning |
120
- |---|---|
121
- | 0 | Evolved OK, or dry-run, or degraded (Darwin absent) |
122
- | 1 | `--alert-on-no-improvement` and champion did not beat parent |
123
- | 2 | Config error or evolution infrastructure failure |
124
- | 99 | Upstream "safety-disqualified" (PROPAGATED, not remapped) |
125
-
126
- ## Graceful degradation (ADR-150 constraint 3 + ADR-153)
127
-
128
- When `@metaharness/darwin` is not installed, the script emits
129
- `{degraded: true, reason: 'metaharness-darwin-not-available', hint: ...}`
130
- and exits 0. ruflo continues to function. CI's
131
- `no-metaharness-smoke.yml`-style job asserts this path.
1
+ ---
2
+ name: harness-evolve
3
+ description: Run `@metaharness/darwin evolve <repo>` to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and promote only measured wins. The model is frozen; the harness evolves. Closes the loop ADR-150 opens (score+genome describe; evolve changes). Degrades gracefully when @metaharness/darwin is absent (ADR-150 + ADR-153 architectural constraints).
4
+ argument-hint: "--repo <path> [--generations 3] [--children 3] [--concurrency 2] [--sandbox real|mock|agent] [--selection pareto|quality-diversity|...] [--mutator deterministic|ruvllm] [--diagnose] [--confirm]"
5
+ allowed-tools: Bash
6
+ ---
7
+
8
+ Surfaces the upstream `metaharness-darwin evolve` CLI as a ruflo skill. The
9
+ **write** layer that pairs with ADR-150's read layer (score / genome /
10
+ mcp-scan / threat-model / oia-audit). Use when you have a harness whose
11
+ readiness scores are flat and you want to discover *which* surface mutation
12
+ moves them — without retraining the foundation model.
13
+
14
+ ## When to use
15
+
16
+ - A `harness-score` result is below target and you don't know which policy
17
+ surface is responsible.
18
+ - You're seeding a harness for a new vertical and want to find a good
19
+ starting configuration empirically rather than hand-tuning.
20
+ - You're comparing your hand-tuned harness against an evolved baseline
21
+ (treat darwin's champion as the strawman).
22
+
23
+ ## When NOT to use
24
+
25
+ - For continuous background optimization. Darwin Mode is human-initiated.
26
+ Wire it into CI for one-shot exploration, not for autonomous self-modification.
27
+ - For ruflo itself in CI. ADR-153 §5 explicitly rejects auto-evolving ruflo
28
+ — the CI gate verifies graceful degradation, not convergence.
29
+
30
+ ## Algorithm
31
+
32
+ Implementation: [`scripts/evolve.mjs`](../../scripts/evolve.mjs).
33
+
34
+ 1. Validate args (`--repo` exists, caps on `--generations` ≤ 50, `--children`
35
+ ≤ 20, `--concurrency` ≤ 8, sandbox/selection/mutator are known values).
36
+ 2. Without `--confirm`: print plan + exit 0 (mirrors `harness-mint` safety
37
+ convention; defense in depth over the upstream `safety.ts` checks).
38
+ 3. With `--confirm`: shell to `npx -y @metaharness/darwin@~0.8.0 metaharness-darwin evolve <repo> ...`
39
+ via the shared `_darwin.mjs` async helper. Per-generation progress is
40
+ forwarded to stderr; final champion JSON is captured from stdout.
41
+ 4. Compute timeout from `generations × children × per-variant` (per-variant
42
+ ≈ 60s real, ≈ 2s mock). Caller may override with `--timeout-ms`.
43
+ 5. Honor upstream exit code 99 — propagate as "safety-disqualified", do not
44
+ remap. This is a designed-in tripwire (a variant tripped `inspectVariant`
45
+ for secrets / shell-out / network / dynamic-eval). See ADR-153 §"Safety model".
46
+ 6. Optional `--alert-on-no-improvement`: exit 1 when champion ≤ parent.
47
+
48
+ ## The seven mutation surfaces
49
+
50
+ | Surface | What it owns |
51
+ |---|---|
52
+ | `planner` | task decomposition / step ordering |
53
+ | `contextBuilder` | what gets fed into the prompt |
54
+ | `reviewer` | self-critique / output verification |
55
+ | `retryPolicy` | when + how to retry on failure |
56
+ | `toolPolicy` | which tools the agent may use, under which conditions |
57
+ | `memoryPolicy` | what to persist, recall, forget |
58
+ | `scorePolicy` | how the agent grades its own output |
59
+
60
+ One mutation per variant. Multi-surface mutations are not allowed (causal
61
+ attribution stays clean).
62
+
63
+ ## Output
64
+
65
+ Reports land under `<repo>/.metaharness/`:
66
+
67
+ ```
68
+ .metaharness/
69
+ archive.json # full lineage tree (sampling next gen draws from this)
70
+ lineage.json # parent→child edges only
71
+ variants/<id>/ # per-variant code (kept for audit)
72
+ runs/<id>/ # per-variant sandbox test output
73
+ reports/winner.json # final champion + score delta vs parent
74
+ ```
75
+
76
+ Skill stdout = JSON `{success, data: {champion, plan, durationMs, improved}}`
77
+ (plus `data.diagnosis` when `--diagnose` is passed — see below).
78
+
79
+ ## Failure diagnosis (`--diagnose`)
80
+
81
+ GEPA's key trick is natural-language failure diagnosis from execution traces
82
+ feeding the next mutation — not just scalar fitness. `--diagnose` adds a
83
+ modest slice of that: after the evolution completes, the losing / failed
84
+ variants' transcripts are run through darwin's GEPA library ops
85
+ (`analyzeTranscript` + `classifyFailure`, via the shared `importGepa`
86
+ resolver in `scripts/_darwin.mjs`) and a `diagnosis` section is appended to
87
+ the emitted JSON:
88
+
89
+ ```json
90
+ "diagnosis": {
91
+ "available": true,
92
+ "scope": "losing-variants",
93
+ "variants": [
94
+ { "id": "g1_v0", "transcripts": 2,
95
+ "failureClasses": { "exploration-loop": 1, "edit-mechanics": 1 },
96
+ "dominantClass": "exploration-loop" }
97
+ ],
98
+ "totals": { "exploration-loop": 1, "edit-mechanics": 1 }
99
+ }
100
+ ```
101
+
102
+ Upstream shape caveats (verified against `@metaharness/darwin@0.8.0`):
103
+
104
+ - `metaharness-darwin evolve --json` prints a TEXT leaderboard — the stdout
105
+ carries no JSON and no transcripts. Per-variant run records live at
106
+ `<repo>/.metaharness/runs/<id>.json`.
107
+ - Those run records hold sandbox exec traces (`{taskId, exitCode, stdout,
108
+ stderr}`), which are NOT GEPA `{actionRaw, obs}` transcripts. Diagnosis
109
+ therefore uses GEPA-shaped transcripts when a run record embeds them
110
+ (agent sandbox / future upstream), falls back to the champion's transcript,
111
+ and otherwise emits `diagnosis: {available: false, reason, traceSummary}`
112
+ where `traceSummary` is a mechanical per-variant tally (tasks / failed /
113
+ timedOut / blockedActions).
114
+ - `--diagnose` NEVER fails the run — any internal error degrades to
115
+ `{available: false, reason: "diagnosis-failed: ..."}`.
116
+
117
+ ## Exit codes
118
+
119
+ | Code | Meaning |
120
+ |---|---|
121
+ | 0 | Evolved OK, or dry-run, or degraded (Darwin absent) |
122
+ | 1 | `--alert-on-no-improvement` and champion did not beat parent |
123
+ | 2 | Config error or evolution infrastructure failure |
124
+ | 99 | Upstream "safety-disqualified" (PROPAGATED, not remapped) |
125
+
126
+ ## Graceful degradation (ADR-150 constraint 3 + ADR-153)
127
+
128
+ When `@metaharness/darwin` is not installed, the script emits
129
+ `{degraded: true, reason: 'metaharness-darwin-not-available', hint: ...}`
130
+ and exits 0. ruflo continues to function. CI's
131
+ `no-metaharness-smoke.yml`-style job asserts this path.
@@ -1,54 +1,54 @@
1
- ---
2
- name: harness-genome
3
- description: 7-section repo readiness report from `metaharness genome <path>`. Returns repo_type / agent_topology / risk_score / mcp_surface / test_confidence / publish_readiness. Pure-read; degrades gracefully (ADR-150).
4
- argument-hint: "[--path .] [--alert-on-risk-above 0.5] [--format table|json]"
5
- allowed-tools: Bash
6
- ---
7
-
8
- Companion to `harness-score`. Where score is a 5-dimension numeric
9
- scorecard, genome is a 7-section categorical/numeric report covering
10
- repo type, agent topology recommendations, risk score (0-1), MCP
11
- surface area, test confidence (0-1), and publish readiness (0-1).
12
-
13
- ## Algorithm
14
-
15
- Implementation: [`scripts/genome.mjs`](../../scripts/genome.mjs).
16
-
17
- 1. Shell out to `npx metaharness genome <path> --json` (60s hard timeout).
18
- 2. Parse the shape: `{ repo_type, agent_topology[], risk_score,
19
- mcp_surface, test_confidence, publish_readiness }`.
20
- 3. If `--alert-on-risk-above N`: exit 1 when `risk_score > N`.
21
- 4. Output JSON (default) or markdown.
22
-
23
- ## Phase-0 baseline (ruflo, measured 2026-06-16)
24
-
25
- ```
26
- {
27
- "repo_type": "node_mcp_ci",
28
- "agent_topology": ["maintainer", "tester", "security", "release"],
29
- "risk_score": 0.27,
30
- "mcp_surface": "remote",
31
- "test_confidence": 0.8,
32
- "publish_readiness": 0.9
33
- }
34
- ```
35
-
36
- Ruflo's `risk_score: 0.27` is low (good). `publish_readiness: 0.9` is
37
- high. The `mcp_surface: "remote"` reflects that ruflo's MCP servers are
38
- hosted, not bundled.
39
-
40
- ## When to use
41
-
42
- - Pre-mint review: "before scaffolding a custom harness from this repo,
43
- should we?" — genome answers it categorically.
44
- - Drift detection: capture genome snapshots over time, diff via
45
- cost-diff-style tooling to spot when `agent_topology` recommendations
46
- drift away from a deliberate architecture choice.
47
- - CI gate: `--alert-on-risk-above 0.5` fails the build when the repo's
48
- risk profile crosses a threshold.
49
-
50
- ## Pairs with
51
-
52
- - `harness-score` — numeric readiness
53
- - `harness-mcp-scan` — static MCP security findings
54
- - `harness-threat-model` — enterprise-review-grade threat model
1
+ ---
2
+ name: harness-genome
3
+ description: 7-section repo readiness report from `metaharness genome <path>`. Returns repo_type / agent_topology / risk_score / mcp_surface / test_confidence / publish_readiness. Pure-read; degrades gracefully (ADR-150).
4
+ argument-hint: "[--path .] [--alert-on-risk-above 0.5] [--format table|json]"
5
+ allowed-tools: Bash
6
+ ---
7
+
8
+ Companion to `harness-score`. Where score is a 5-dimension numeric
9
+ scorecard, genome is a 7-section categorical/numeric report covering
10
+ repo type, agent topology recommendations, risk score (0-1), MCP
11
+ surface area, test confidence (0-1), and publish readiness (0-1).
12
+
13
+ ## Algorithm
14
+
15
+ Implementation: [`scripts/genome.mjs`](../../scripts/genome.mjs).
16
+
17
+ 1. Shell out to `npx metaharness genome <path> --json` (60s hard timeout).
18
+ 2. Parse the shape: `{ repo_type, agent_topology[], risk_score,
19
+ mcp_surface, test_confidence, publish_readiness }`.
20
+ 3. If `--alert-on-risk-above N`: exit 1 when `risk_score > N`.
21
+ 4. Output JSON (default) or markdown.
22
+
23
+ ## Phase-0 baseline (ruflo, measured 2026-06-16)
24
+
25
+ ```
26
+ {
27
+ "repo_type": "node_mcp_ci",
28
+ "agent_topology": ["maintainer", "tester", "security", "release"],
29
+ "risk_score": 0.27,
30
+ "mcp_surface": "remote",
31
+ "test_confidence": 0.8,
32
+ "publish_readiness": 0.9
33
+ }
34
+ ```
35
+
36
+ Ruflo's `risk_score: 0.27` is low (good). `publish_readiness: 0.9` is
37
+ high. The `mcp_surface: "remote"` reflects that ruflo's MCP servers are
38
+ hosted, not bundled.
39
+
40
+ ## When to use
41
+
42
+ - Pre-mint review: "before scaffolding a custom harness from this repo,
43
+ should we?" — genome answers it categorically.
44
+ - Drift detection: capture genome snapshots over time, diff via
45
+ cost-diff-style tooling to spot when `agent_topology` recommendations
46
+ drift away from a deliberate architecture choice.
47
+ - CI gate: `--alert-on-risk-above 0.5` fails the build when the repo's
48
+ risk profile crosses a threshold.
49
+
50
+ ## Pairs with
51
+
52
+ - `harness-score` — numeric readiness
53
+ - `harness-mcp-scan` — static MCP security findings
54
+ - `harness-threat-model` — enterprise-review-grade threat model