@claude-flow/cli 3.32.9 → 3.32.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/agents/analysis/analyze-code-quality.md +178 -178
- package/.claude/agents/analysis/code-analyzer.md +209 -209
- package/.claude/agents/analysis/code-review/analyze-code-quality.md +178 -178
- package/.claude/agents/architecture/arch-system-design.md +156 -156
- package/.claude/agents/architecture/system-design/arch-system-design.md +154 -154
- package/.claude/agents/browser/browser-agent.yaml +182 -182
- package/.claude/agents/consensus/byzantine-coordinator.md +62 -62
- package/.claude/agents/consensus/crdt-synchronizer.md +996 -996
- package/.claude/agents/consensus/gossip-coordinator.md +62 -62
- package/.claude/agents/consensus/performance-benchmarker.md +850 -850
- package/.claude/agents/consensus/quorum-manager.md +822 -822
- package/.claude/agents/consensus/raft-manager.md +62 -62
- package/.claude/agents/consensus/security-manager.md +621 -621
- package/.claude/agents/core/planner.md +374 -374
- package/.claude/agents/custom/test-long-runner.md +44 -44
- package/.claude/agents/data/data-ml-model.md +444 -444
- package/.claude/agents/data/ml/data-ml-model.md +192 -192
- package/.claude/agents/development/backend/dev-backend-api.md +141 -141
- package/.claude/agents/development/dev-backend-api.md +344 -344
- package/.claude/agents/devops/ci-cd/ops-cicd-github.md +163 -163
- package/.claude/agents/devops/ops-cicd-github.md +164 -164
- package/.claude/agents/documentation/api-docs/docs-api-openapi.md +173 -173
- package/.claude/agents/documentation/docs-api-openapi.md +354 -354
- package/.claude/agents/flow-nexus/app-store.md +87 -87
- package/.claude/agents/flow-nexus/authentication.md +68 -68
- package/.claude/agents/flow-nexus/challenges.md +80 -80
- package/.claude/agents/flow-nexus/neural-network.md +87 -87
- package/.claude/agents/flow-nexus/payments.md +82 -82
- package/.claude/agents/flow-nexus/sandbox.md +75 -75
- package/.claude/agents/flow-nexus/swarm.md +75 -75
- package/.claude/agents/flow-nexus/user-tools.md +95 -95
- package/.claude/agents/flow-nexus/workflow.md +83 -83
- package/.claude/agents/github/code-review-swarm.md +377 -377
- package/.claude/agents/github/github-modes.md +172 -172
- package/.claude/agents/github/issue-tracker.md +575 -575
- package/.claude/agents/github/multi-repo-swarm.md +552 -552
- package/.claude/agents/github/pr-manager.md +437 -437
- package/.claude/agents/github/project-board-sync.md +508 -508
- package/.claude/agents/github/release-manager.md +604 -604
- package/.claude/agents/github/release-swarm.md +582 -582
- package/.claude/agents/github/repo-architect.md +397 -397
- package/.claude/agents/github/swarm-issue.md +572 -572
- package/.claude/agents/github/swarm-pr.md +427 -427
- package/.claude/agents/github/sync-coordinator.md +451 -451
- package/.claude/agents/github/workflow-automation.md +902 -902
- package/.claude/agents/goal/agent.md +815 -815
- package/.claude/agents/optimization/benchmark-suite.md +664 -664
- package/.claude/agents/optimization/load-balancer.md +430 -430
- package/.claude/agents/optimization/performance-monitor.md +671 -671
- package/.claude/agents/optimization/resource-allocator.md +673 -673
- package/.claude/agents/optimization/topology-optimizer.md +807 -807
- package/.claude/agents/payments/agentic-payments.md +126 -126
- package/.claude/agents/sona/sona-learning-optimizer.md +74 -74
- package/.claude/agents/sparc/architecture.md +698 -698
- package/.claude/agents/sparc/pseudocode.md +519 -519
- package/.claude/agents/sparc/refinement.md +801 -801
- package/.claude/agents/sparc/specification.md +477 -477
- package/.claude/agents/specialized/mobile/spec-mobile-react-native.md +224 -224
- package/.claude/agents/specialized/spec-mobile-react-native.md +226 -226
- package/.claude/agents/sublinear/consensus-coordinator.md +337 -337
- package/.claude/agents/sublinear/matrix-optimizer.md +184 -184
- package/.claude/agents/sublinear/pagerank-analyzer.md +298 -298
- package/.claude/agents/sublinear/performance-optimizer.md +367 -367
- package/.claude/agents/sublinear/trading-predictor.md +245 -245
- package/.claude/agents/swarm/adaptive-coordinator.md +1126 -1126
- package/.claude/agents/swarm/hierarchical-coordinator.md +709 -709
- package/.claude/agents/swarm/mesh-coordinator.md +962 -962
- package/.claude/agents/templates/automation-smart-agent.md +204 -204
- package/.claude/agents/templates/base-template-generator.md +289 -289
- package/.claude/agents/templates/coordinator-swarm-init.md +89 -89
- package/.claude/agents/templates/github-pr-manager.md +176 -176
- package/.claude/agents/templates/implementer-sparc-coder.md +258 -258
- package/.claude/agents/templates/memory-coordinator.md +186 -186
- package/.claude/agents/templates/orchestrator-task.md +138 -138
- package/.claude/agents/templates/performance-analyzer.md +198 -198
- package/.claude/agents/templates/sparc-coordinator.md +513 -513
- package/.claude/agents/testing/production-validator.md +394 -394
- package/.claude/agents/testing/tdd-london-swarm.md +243 -243
- package/.claude/agents/v3/aidefence-guardian.md +282 -282
- package/.claude/agents/v3/claims-authorizer.md +208 -208
- package/.claude/agents/v3/collective-intelligence-coordinator.md +993 -993
- package/.claude/agents/v3/ddd-domain-expert.md +220 -220
- package/.claude/agents/v3/injection-analyst.md +236 -236
- package/.claude/agents/v3/performance-engineer.md +1233 -1233
- package/.claude/agents/v3/pii-detector.md +151 -151
- package/.claude/agents/v3/reasoningbank-learner.md +213 -213
- package/.claude/agents/v3/security-architect-aidefence.md +410 -410
- package/.claude/agents/v3/security-architect.md +867 -867
- package/.claude/agents/v3/swarm-memory-manager.md +157 -157
- package/.claude/agents/v3/v3-integration-architect.md +205 -205
- package/.claude/commands/agents/README.md +50 -50
- package/.claude/commands/agents/agent-capabilities.md +140 -140
- package/.claude/commands/agents/agent-coordination.md +28 -28
- package/.claude/commands/agents/agent-spawning.md +28 -28
- package/.claude/commands/agents/agent-types.md +216 -216
- package/.claude/commands/agents/health.md +139 -139
- package/.claude/commands/agents/list.md +100 -100
- package/.claude/commands/agents/logs.md +130 -130
- package/.claude/commands/agents/metrics.md +122 -122
- package/.claude/commands/agents/pool.md +127 -127
- package/.claude/commands/agents/spawn.md +140 -140
- package/.claude/commands/agents/status.md +115 -115
- package/.claude/commands/agents/stop.md +102 -102
- package/.claude/commands/analysis/COMMAND_COMPLIANCE_REPORT.md +53 -53
- package/.claude/commands/analysis/README.md +9 -9
- package/.claude/commands/analysis/bottleneck-detect.md +162 -162
- package/.claude/commands/analysis/performance-bottlenecks.md +58 -58
- package/.claude/commands/analysis/performance-report.md +25 -25
- package/.claude/commands/analysis/token-efficiency.md +44 -44
- package/.claude/commands/analysis/token-usage.md +25 -25
- package/.claude/commands/automation/README.md +9 -9
- package/.claude/commands/automation/auto-agent.md +122 -122
- package/.claude/commands/automation/self-healing.md +105 -105
- package/.claude/commands/automation/session-memory.md +89 -89
- package/.claude/commands/automation/smart-agents.md +72 -72
- package/.claude/commands/automation/smart-spawn.md +25 -25
- package/.claude/commands/automation/workflow-select.md +25 -25
- package/.claude/commands/claude-flow-help.md +103 -103
- package/.claude/commands/claude-flow-memory.md +107 -107
- package/.claude/commands/claude-flow-swarm.md +205 -205
- package/.claude/commands/coordination/README.md +9 -9
- package/.claude/commands/coordination/agent-spawn.md +25 -25
- package/.claude/commands/coordination/init.md +44 -44
- package/.claude/commands/coordination/orchestrate.md +43 -43
- package/.claude/commands/coordination/spawn.md +45 -45
- package/.claude/commands/coordination/swarm-init.md +85 -85
- package/.claude/commands/coordination/task-orchestrate.md +25 -25
- package/.claude/commands/github/README.md +11 -11
- package/.claude/commands/github/code-review-swarm.md +513 -513
- package/.claude/commands/github/code-review.md +25 -25
- package/.claude/commands/github/github-modes.md +146 -146
- package/.claude/commands/github/github-swarm.md +121 -121
- package/.claude/commands/github/issue-tracker.md +291 -291
- package/.claude/commands/github/issue-triage.md +25 -25
- package/.claude/commands/github/multi-repo-swarm.md +518 -518
- package/.claude/commands/github/pr-enhance.md +26 -26
- package/.claude/commands/github/pr-manager.md +169 -169
- package/.claude/commands/github/project-board-sync.md +470 -470
- package/.claude/commands/github/release-manager.md +339 -339
- package/.claude/commands/github/release-swarm.md +543 -543
- package/.claude/commands/github/repo-analyze.md +25 -25
- package/.claude/commands/github/repo-architect.md +366 -366
- package/.claude/commands/github/swarm-issue.md +484 -484
- package/.claude/commands/github/swarm-pr.md +287 -287
- package/.claude/commands/github/sync-coordinator.md +302 -302
- package/.claude/commands/github/workflow-automation.md +441 -441
- package/.claude/commands/hive-mind/README.md +17 -17
- package/.claude/commands/hive-mind/hive-mind-consensus.md +8 -8
- package/.claude/commands/hive-mind/hive-mind-init.md +18 -18
- package/.claude/commands/hive-mind/hive-mind-memory.md +8 -8
- package/.claude/commands/hive-mind/hive-mind-metrics.md +8 -8
- package/.claude/commands/hive-mind/hive-mind-resume.md +8 -8
- package/.claude/commands/hive-mind/hive-mind-sessions.md +8 -8
- package/.claude/commands/hive-mind/hive-mind-spawn.md +21 -21
- package/.claude/commands/hive-mind/hive-mind-status.md +8 -8
- package/.claude/commands/hive-mind/hive-mind-stop.md +8 -8
- package/.claude/commands/hive-mind/hive-mind-wizard.md +8 -8
- package/.claude/commands/hive-mind/hive-mind.md +27 -27
- package/.claude/commands/hooks/README.md +11 -11
- package/.claude/commands/hooks/overview.md +57 -57
- package/.claude/commands/hooks/post-edit.md +117 -117
- package/.claude/commands/hooks/post-task.md +112 -112
- package/.claude/commands/hooks/pre-edit.md +113 -113
- package/.claude/commands/hooks/pre-task.md +111 -111
- package/.claude/commands/hooks/session-end.md +118 -118
- package/.claude/commands/hooks/setup.md +102 -102
- package/.claude/commands/memory/README.md +9 -9
- package/.claude/commands/memory/memory-persist.md +25 -25
- package/.claude/commands/memory/memory-search.md +25 -25
- package/.claude/commands/memory/memory-usage.md +25 -25
- package/.claude/commands/memory/neural.md +47 -47
- package/.claude/commands/monitoring/README.md +9 -9
- package/.claude/commands/monitoring/agent-metrics.md +25 -25
- package/.claude/commands/monitoring/agents.md +44 -44
- package/.claude/commands/monitoring/real-time-view.md +25 -25
- package/.claude/commands/monitoring/status.md +46 -46
- package/.claude/commands/monitoring/swarm-monitor.md +25 -25
- package/.claude/commands/optimization/README.md +9 -9
- package/.claude/commands/optimization/auto-topology.md +61 -61
- package/.claude/commands/optimization/cache-manage.md +25 -25
- package/.claude/commands/optimization/parallel-execute.md +25 -25
- package/.claude/commands/optimization/parallel-execution.md +49 -49
- package/.claude/commands/optimization/topology-optimize.md +25 -25
- package/.claude/commands/pair/README.md +260 -260
- package/.claude/commands/pair/commands.md +545 -545
- package/.claude/commands/pair/config.md +509 -509
- package/.claude/commands/pair/examples.md +511 -511
- package/.claude/commands/pair/modes.md +347 -347
- package/.claude/commands/pair/session.md +406 -406
- package/.claude/commands/pair/start.md +208 -208
- package/.claude/commands/sparc/analyzer.md +51 -51
- package/.claude/commands/sparc/architect.md +53 -53
- package/.claude/commands/sparc/ask.md +97 -97
- package/.claude/commands/sparc/batch-executor.md +54 -54
- package/.claude/commands/sparc/code.md +89 -89
- package/.claude/commands/sparc/coder.md +54 -54
- package/.claude/commands/sparc/debug.md +83 -83
- package/.claude/commands/sparc/debugger.md +54 -54
- package/.claude/commands/sparc/designer.md +53 -53
- package/.claude/commands/sparc/devops.md +109 -109
- package/.claude/commands/sparc/docs-writer.md +80 -80
- package/.claude/commands/sparc/documenter.md +54 -54
- package/.claude/commands/sparc/innovator.md +54 -54
- package/.claude/commands/sparc/integration.md +83 -83
- package/.claude/commands/sparc/mcp.md +117 -117
- package/.claude/commands/sparc/memory-manager.md +54 -54
- package/.claude/commands/sparc/optimizer.md +54 -54
- package/.claude/commands/sparc/orchestrator.md +131 -131
- package/.claude/commands/sparc/post-deployment-monitoring-mode.md +83 -83
- package/.claude/commands/sparc/refinement-optimization-mode.md +83 -83
- package/.claude/commands/sparc/researcher.md +54 -54
- package/.claude/commands/sparc/reviewer.md +54 -54
- package/.claude/commands/sparc/security-review.md +80 -80
- package/.claude/commands/sparc/sparc-modes.md +174 -174
- package/.claude/commands/sparc/sparc.md +111 -111
- package/.claude/commands/sparc/spec-pseudocode.md +80 -80
- package/.claude/commands/sparc/supabase-admin.md +348 -348
- package/.claude/commands/sparc/swarm-coordinator.md +54 -54
- package/.claude/commands/sparc/tdd.md +54 -54
- package/.claude/commands/sparc/tester.md +54 -54
- package/.claude/commands/sparc/tutorial.md +79 -79
- package/.claude/commands/sparc/workflow-manager.md +54 -54
- package/.claude/commands/sparc.md +166 -166
- package/.claude/commands/stream-chain/pipeline.md +120 -120
- package/.claude/commands/stream-chain/run.md +69 -69
- package/.claude/commands/swarm/README.md +15 -15
- package/.claude/commands/swarm/analysis.md +95 -95
- package/.claude/commands/swarm/development.md +96 -96
- package/.claude/commands/swarm/examples.md +168 -168
- package/.claude/commands/swarm/maintenance.md +102 -102
- package/.claude/commands/swarm/optimization.md +117 -117
- package/.claude/commands/swarm/research.md +136 -136
- package/.claude/commands/swarm/swarm-analysis.md +8 -8
- package/.claude/commands/swarm/swarm-background.md +8 -8
- package/.claude/commands/swarm/swarm-init.md +19 -19
- package/.claude/commands/swarm/swarm-modes.md +8 -8
- package/.claude/commands/swarm/swarm-monitor.md +8 -8
- package/.claude/commands/swarm/swarm-spawn.md +19 -19
- package/.claude/commands/swarm/swarm-status.md +8 -8
- package/.claude/commands/swarm/swarm-strategies.md +8 -8
- package/.claude/commands/swarm/swarm.md +87 -87
- package/.claude/commands/swarm/testing.md +131 -131
- package/.claude/commands/training/README.md +9 -9
- package/.claude/commands/training/model-update.md +25 -25
- package/.claude/commands/training/neural-patterns.md +107 -107
- package/.claude/commands/training/neural-train.md +75 -75
- package/.claude/commands/training/pattern-learn.md +25 -25
- package/.claude/commands/training/specialization.md +62 -62
- package/.claude/commands/truth/start.md +142 -142
- package/.claude/commands/verify/check.md +49 -49
- package/.claude/commands/verify/start.md +127 -127
- package/.claude/commands/workflows/README.md +9 -9
- package/.claude/commands/workflows/development.md +77 -77
- package/.claude/commands/workflows/research.md +62 -62
- package/.claude/commands/workflows/workflow-create.md +25 -25
- package/.claude/commands/workflows/workflow-execute.md +25 -25
- package/.claude/commands/workflows/workflow-export.md +25 -25
- package/.claude/eval/human-relevance-frozen-v1.json +17 -17
- package/.claude/evolve-proof/generation-0.json +211 -211
- package/.claude/evolve-proof/real-generation-0.json +406 -406
- package/.claude/evolve-proof/real-generation-1.json +406 -406
- package/.claude/helpers/README.md +96 -96
- package/.claude/helpers/adr-compliance.sh +186 -186
- package/.claude/helpers/auto-commit.sh +178 -178
- package/.claude/helpers/auto-memory-hook.mjs +430 -430
- package/.claude/helpers/checkpoint-manager.sh +251 -251
- package/.claude/helpers/daemon-manager.sh +252 -252
- package/.claude/helpers/ddd-tracker.sh +144 -144
- package/.claude/helpers/github-safe.js +156 -156
- package/.claude/helpers/github-setup.sh +45 -45
- package/.claude/helpers/guidance-hook.sh +13 -13
- package/.claude/helpers/guidance-hooks.sh +102 -102
- package/.claude/helpers/health-monitor.sh +108 -108
- package/.claude/helpers/helpers.manifest.json +6 -6
- package/.claude/helpers/hook-handler.cjs +565 -565
- package/.claude/helpers/intelligence.cjs +1058 -1058
- package/.claude/helpers/learning-hooks.sh +329 -329
- package/.claude/helpers/learning-optimizer.sh +127 -127
- package/.claude/helpers/learning-service.mjs +1144 -1144
- package/.claude/helpers/memory.js +83 -83
- package/.claude/helpers/metrics-db.mjs +503 -503
- package/.claude/helpers/pattern-consolidator.sh +86 -86
- package/.claude/helpers/perf-worker.sh +160 -160
- package/.claude/helpers/post-commit +16 -16
- package/.claude/helpers/pre-commit +26 -26
- package/.claude/helpers/quick-start.sh +19 -19
- package/.claude/helpers/router.js +105 -105
- package/.claude/helpers/security-scanner.sh +127 -127
- package/.claude/helpers/session.js +157 -157
- package/.claude/helpers/setup-mcp.sh +18 -18
- package/.claude/helpers/standard-checkpoint-hooks.sh +189 -189
- package/.claude/helpers/statusline-hook.sh +21 -21
- package/.claude/helpers/statusline.cjs +1060 -1060
- package/.claude/helpers/statusline.js +340 -340
- package/.claude/helpers/swarm-comms.sh +353 -353
- package/.claude/helpers/swarm-hooks.sh +761 -761
- package/.claude/helpers/swarm-monitor.sh +210 -210
- package/.claude/helpers/sync-v3-metrics.sh +245 -245
- package/.claude/helpers/update-v3-progress.sh +165 -165
- package/.claude/helpers/v3-quick-status.sh +57 -57
- package/.claude/helpers/v3.sh +110 -110
- package/.claude/helpers/validate-v3-config.sh +215 -215
- package/.claude/helpers/worker-manager.sh +170 -170
- package/.claude/proven-config.manifest.json +37 -37
- package/.claude/proven-config.signed.json +41 -41
- package/.claude/settings.json +182 -182
- package/.claude/skills/agentdb-advanced/SKILL.md +550 -550
- package/.claude/skills/agentdb-learning/SKILL.md +545 -545
- package/.claude/skills/agentdb-memory-patterns/SKILL.md +339 -339
- package/.claude/skills/agentdb-optimization/SKILL.md +509 -509
- package/.claude/skills/agentdb-vector-search/SKILL.md +339 -339
- package/.claude/skills/browser/SKILL.md +204 -204
- package/.claude/skills/dual-mode/README.md +71 -71
- package/.claude/skills/dual-mode/dual-collect.md +103 -103
- package/.claude/skills/dual-mode/dual-coordinate.md +85 -85
- package/.claude/skills/dual-mode/dual-spawn.md +81 -81
- package/.claude/skills/flow-nexus-neural/SKILL.md +727 -727
- package/.claude/skills/flow-nexus-platform/SKILL.md +1154 -1154
- package/.claude/skills/flow-nexus-swarm/SKILL.md +604 -604
- package/.claude/skills/github-code-review/SKILL.md +1125 -1125
- package/.claude/skills/github-multi-repo/SKILL.md +862 -862
- package/.claude/skills/github-project-management/SKILL.md +1262 -1262
- package/.claude/skills/github-release-management/SKILL.md +1064 -1064
- package/.claude/skills/github-workflow-automation/SKILL.md +1047 -1047
- package/.claude/skills/hooks-automation/SKILL.md +1201 -1201
- package/.claude/skills/pair-programming/SKILL.md +1202 -1202
- package/.claude/skills/reasoningbank-agentdb/SKILL.md +446 -446
- package/.claude/skills/reasoningbank-intelligence/SKILL.md +201 -201
- package/.claude/skills/skill-builder/SKILL.md +910 -910
- package/.claude/skills/sparc-methodology/SKILL.md +1106 -1106
- package/.claude/skills/stream-chain/SKILL.md +560 -560
- package/.claude/skills/swarm-advanced/SKILL.md +970 -970
- package/.claude/skills/swarm-orchestration/SKILL.md +179 -179
- package/.claude/skills/v3-cli-modernization/SKILL.md +871 -871
- package/.claude/skills/v3-core-implementation/SKILL.md +796 -796
- package/.claude/skills/v3-ddd-architecture/SKILL.md +441 -441
- package/.claude/skills/v3-integration-deep/SKILL.md +240 -240
- package/.claude/skills/v3-mcp-optimization/SKILL.md +776 -776
- package/.claude/skills/v3-memory-unification/SKILL.md +173 -173
- package/.claude/skills/v3-performance-optimization/SKILL.md +389 -389
- package/.claude/skills/v3-security-overhaul/SKILL.md +81 -81
- package/.claude/skills/v3-swarm-coordination/SKILL.md +339 -339
- package/.claude/skills/verification-quality/SKILL.md +691 -691
- package/README.md +419 -419
- package/bin/cli.js +314 -314
- package/bin/mcp-server.js +224 -224
- package/bin/preinstall.cjs +2 -2
- package/catalog-manifest.json +2 -2
- package/dist/src/autopilot-state.js +24 -7
- package/dist/src/benchmarks/gaia-critic.js +24 -24
- package/dist/src/business-pods/bbs-budget-tracker.js +53 -53
- package/dist/src/commands/completions.js +409 -409
- package/dist/src/commands/daemon.js +44 -44
- package/dist/src/commands/embeddings.js +26 -26
- package/dist/src/commands/hive-mind.js +97 -97
- package/dist/src/commands/hooks.js +31 -10
- package/dist/src/commands/init.js +202 -34
- package/dist/src/commands/memory.js +12 -1
- package/dist/src/commands/ruvector/backup.js +23 -23
- package/dist/src/commands/ruvector/benchmark.js +31 -31
- package/dist/src/commands/ruvector/import.js +14 -14
- package/dist/src/commands/ruvector/init.js +115 -115
- package/dist/src/commands/ruvector/migrate.js +99 -99
- package/dist/src/commands/ruvector/optimize.js +51 -51
- package/dist/src/commands/ruvector/setup.js +624 -624
- package/dist/src/commands/ruvector/status.js +38 -38
- package/dist/src/config/proven-config.js +2 -2
- package/dist/src/funnel/disclosure.js +13 -2
- package/dist/src/funnel/messages.d.ts +12 -10
- package/dist/src/funnel/messages.js +83 -11
- package/dist/src/init/claudemd-generator.js +231 -231
- package/dist/src/init/executor.js +453 -453
- package/dist/src/init/helper-signing.js +2 -2
- package/dist/src/init/helpers-generator.js +751 -751
- package/dist/src/init/statusline-generator.js +24 -24
- package/dist/src/mcp-tools/agentdb-tools.js +15 -15
- package/dist/src/mcp-tools/browser-intent-tools.js +19 -19
- package/dist/src/mcp-tools/browser-tools.js +8 -0
- package/dist/src/mcp-tools/hooks-tools.js +21 -0
- package/dist/src/mcp-tools/memory-tools.js +4 -3
- package/dist/src/memory/graph-edge-writer.js +22 -22
- package/dist/src/memory/memory-bridge.js +192 -123
- package/dist/src/memory/memory-initializer.js +407 -407
- package/dist/src/memory/rabitq-index.js +5 -5
- package/dist/src/parser.js +25 -9
- package/dist/src/proxy/verify.js +2 -2
- package/dist/src/runtime/headless.js +28 -28
- package/dist/src/services/distill-tuning.js +7 -7
- package/dist/src/services/headless-worker-executor.js +84 -84
- package/dist/src/services/memory-distillation.js +4 -4
- package/dist/src/services/worker-daemon.js +7 -4
- package/dist/src/transfer/deploy-seraphine.js +23 -23
- package/package.json +137 -137
- package/plugins/ruflo-metaharness/.claude-plugin/plugin.json +32 -32
- package/plugins/ruflo-metaharness/README.md +72 -72
- package/plugins/ruflo-metaharness/agents/metaharness-architect.md +58 -58
- package/plugins/ruflo-metaharness/commands/ruflo-metaharness.md +48 -48
- package/plugins/ruflo-metaharness/scripts/_darwin.mjs +210 -210
- package/plugins/ruflo-metaharness/scripts/_harness.mjs +330 -330
- package/plugins/ruflo-metaharness/scripts/_invoke.mjs +231 -231
- package/plugins/ruflo-metaharness/scripts/_redblue.mjs +143 -143
- package/plugins/ruflo-metaharness/scripts/_similarity.mjs +161 -161
- package/plugins/ruflo-metaharness/scripts/_spike-similarity.mjs +223 -223
- package/plugins/ruflo-metaharness/scripts/audit-list.mjs +158 -158
- package/plugins/ruflo-metaharness/scripts/audit-trend.mjs +272 -272
- package/plugins/ruflo-metaharness/scripts/bench-parse-mcp-scan.mjs +146 -146
- package/plugins/ruflo-metaharness/scripts/bench-recordpair-overhead.mjs +186 -186
- package/plugins/ruflo-metaharness/scripts/bench-similarity.mjs +177 -177
- package/plugins/ruflo-metaharness/scripts/bench.mjs +95 -95
- package/plugins/ruflo-metaharness/scripts/drift-from-history.mjs +363 -363
- package/plugins/ruflo-metaharness/scripts/evolve.mjs +404 -404
- package/plugins/ruflo-metaharness/scripts/genome.mjs +80 -80
- package/plugins/ruflo-metaharness/scripts/gepa.mjs +153 -153
- package/plugins/ruflo-metaharness/scripts/learn.mjs +127 -127
- package/plugins/ruflo-metaharness/scripts/mcp-scan.mjs +111 -111
- package/plugins/ruflo-metaharness/scripts/mint.mjs +126 -126
- package/plugins/ruflo-metaharness/scripts/oia-audit.mjs +228 -228
- package/plugins/ruflo-metaharness/scripts/redblue.mjs +286 -286
- package/plugins/ruflo-metaharness/scripts/router-parallel-analyze.mjs +250 -250
- package/plugins/ruflo-metaharness/scripts/score.mjs +92 -92
- package/plugins/ruflo-metaharness/scripts/security-bench.mjs +174 -174
- package/plugins/ruflo-metaharness/scripts/similarity.mjs +158 -158
- package/plugins/ruflo-metaharness/scripts/smoke.sh +2356 -2356
- package/plugins/ruflo-metaharness/scripts/test-graceful-degradation.mjs +165 -165
- package/plugins/ruflo-metaharness/scripts/test-mcp-tools.mjs +472 -472
- package/plugins/ruflo-metaharness/scripts/test-parallel-pipeline.mjs +204 -204
- package/plugins/ruflo-metaharness/scripts/test-pipeline-roundtrip.mjs +586 -586
- package/plugins/ruflo-metaharness/scripts/test-similarity.mjs +334 -334
- package/plugins/ruflo-metaharness/scripts/test-with-openrouter.mjs +229 -229
- package/plugins/ruflo-metaharness/scripts/threat-model.mjs +59 -59
- package/plugins/ruflo-metaharness/skills/harness-bench/SKILL.md +64 -64
- package/plugins/ruflo-metaharness/skills/harness-drift-from-history/SKILL.md +65 -65
- package/plugins/ruflo-metaharness/skills/harness-evolve/SKILL.md +131 -131
- package/plugins/ruflo-metaharness/skills/harness-genome/SKILL.md +54 -54
- package/plugins/ruflo-metaharness/skills/harness-gepa/SKILL.md +65 -65
- package/plugins/ruflo-metaharness/skills/harness-learn/SKILL.md +65 -65
- package/plugins/ruflo-metaharness/skills/harness-mcp-scan/SKILL.md +49 -49
- package/plugins/ruflo-metaharness/skills/harness-mint/SKILL.md +72 -72
- package/plugins/ruflo-metaharness/skills/harness-oia-audit/SKILL.md +79 -79
- package/plugins/ruflo-metaharness/skills/harness-score/SKILL.md +66 -66
- package/plugins/ruflo-metaharness/skills/harness-security-bench/SKILL.md +101 -101
- package/plugins/ruflo-metaharness/skills/harness-similarity/SKILL.md +67 -67
- package/plugins/ruflo-metaharness/skills/harness-threat-model/SKILL.md +41 -41
- package/scripts/postinstall.cjs +153 -153
|
@@ -1,64 +1,64 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: harness-bench
|
|
3
|
-
description: Manage `@metaharness/darwin` bench suites — `bench create <repo>` scaffolds a JSON suite from a repo's test corpus; `bench verify <suite.json>` checks suite well-formedness. Bench suites are the fixed evaluation corpora that `harness-evolve --bench <suite.json>` scores variants against, decoupling evolution from the repo's natural tests. Degrades gracefully when @metaharness/darwin is absent.
|
|
4
|
-
argument-hint: "--op create --repo <path> [--out <path>] | --op verify --suite <path>"
|
|
5
|
-
allowed-tools: Bash
|
|
6
|
-
---
|
|
7
|
-
|
|
8
|
-
Surfaces `metaharness-darwin bench <create|verify>` — the supporting verb
|
|
9
|
-
for `harness-evolve --bench`. Use when you want evolution scored against a
|
|
10
|
-
fixed corpus (independent of `npm test`) so champion fitness is comparable
|
|
11
|
-
across commits or across forks of the same harness.
|
|
12
|
-
|
|
13
|
-
## When to use
|
|
14
|
-
|
|
15
|
-
- Setting up a new evolution pipeline for a repo whose `npm test` is
|
|
16
|
-
flaky, slow, or undersized — scaffold a deterministic bench suite once,
|
|
17
|
-
then evolve against it repeatedly.
|
|
18
|
-
- CI: `bench verify` the checked-in suite on every PR that touches it
|
|
19
|
-
(cheap; ~5s).
|
|
20
|
-
- Forking a harness to a new domain: copy and edit the suite to retarget
|
|
21
|
-
the evaluation without losing comparability to the parent.
|
|
22
|
-
|
|
23
|
-
## Algorithm
|
|
24
|
-
|
|
25
|
-
Implementation: [`scripts/bench.mjs`](../../scripts/bench.mjs).
|
|
26
|
-
|
|
27
|
-
### `--op create`
|
|
28
|
-
1. Resolve `--repo` path; reject if missing.
|
|
29
|
-
2. Shell to `metaharness-darwin bench create <repo> [--out <suite.json>]`.
|
|
30
|
-
3. Default output path: `<repo>/.metaharness/bench/suite.json` (chosen by upstream).
|
|
31
|
-
4. Suite shape (per upstream): array of `{ input, expectedOutput, weight }` tasks
|
|
32
|
-
derived from existing test cases.
|
|
33
|
-
|
|
34
|
-
### `--op verify`
|
|
35
|
-
1. Resolve `--suite` path; reject if missing.
|
|
36
|
-
2. Shell to `metaharness-darwin bench verify <suite.json>`.
|
|
37
|
-
3. Exit 1 if any task malformed (upstream's signal).
|
|
38
|
-
|
|
39
|
-
## Output shape
|
|
40
|
-
|
|
41
|
-
```json
|
|
42
|
-
{
|
|
43
|
-
"success": true,
|
|
44
|
-
"data": {
|
|
45
|
-
"op": "verify",
|
|
46
|
-
"taskCount": 42,
|
|
47
|
-
"wellFormed": true,
|
|
48
|
-
"durationMs": 870
|
|
49
|
-
}
|
|
50
|
-
}
|
|
51
|
-
```
|
|
52
|
-
|
|
53
|
-
## Exit codes
|
|
54
|
-
|
|
55
|
-
| Code | Meaning |
|
|
56
|
-
|---|---|
|
|
57
|
-
| 0 | OK (or degraded — Darwin absent) |
|
|
58
|
-
| 1 | `--op verify` and suite malformed |
|
|
59
|
-
| 2 | Config error or upstream invocation failure |
|
|
60
|
-
|
|
61
|
-
## Graceful degradation
|
|
62
|
-
|
|
63
|
-
When `@metaharness/darwin` is absent, emits the standard `{degraded: true,
|
|
64
|
-
reason: 'metaharness-darwin-not-available'}` payload and exits 0.
|
|
1
|
+
---
|
|
2
|
+
name: harness-bench
|
|
3
|
+
description: Manage `@metaharness/darwin` bench suites — `bench create <repo>` scaffolds a JSON suite from a repo's test corpus; `bench verify <suite.json>` checks suite well-formedness. Bench suites are the fixed evaluation corpora that `harness-evolve --bench <suite.json>` scores variants against, decoupling evolution from the repo's natural tests. Degrades gracefully when @metaharness/darwin is absent.
|
|
4
|
+
argument-hint: "--op create --repo <path> [--out <path>] | --op verify --suite <path>"
|
|
5
|
+
allowed-tools: Bash
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
Surfaces `metaharness-darwin bench <create|verify>` — the supporting verb
|
|
9
|
+
for `harness-evolve --bench`. Use when you want evolution scored against a
|
|
10
|
+
fixed corpus (independent of `npm test`) so champion fitness is comparable
|
|
11
|
+
across commits or across forks of the same harness.
|
|
12
|
+
|
|
13
|
+
## When to use
|
|
14
|
+
|
|
15
|
+
- Setting up a new evolution pipeline for a repo whose `npm test` is
|
|
16
|
+
flaky, slow, or undersized — scaffold a deterministic bench suite once,
|
|
17
|
+
then evolve against it repeatedly.
|
|
18
|
+
- CI: `bench verify` the checked-in suite on every PR that touches it
|
|
19
|
+
(cheap; ~5s).
|
|
20
|
+
- Forking a harness to a new domain: copy and edit the suite to retarget
|
|
21
|
+
the evaluation without losing comparability to the parent.
|
|
22
|
+
|
|
23
|
+
## Algorithm
|
|
24
|
+
|
|
25
|
+
Implementation: [`scripts/bench.mjs`](../../scripts/bench.mjs).
|
|
26
|
+
|
|
27
|
+
### `--op create`
|
|
28
|
+
1. Resolve `--repo` path; reject if missing.
|
|
29
|
+
2. Shell to `metaharness-darwin bench create <repo> [--out <suite.json>]`.
|
|
30
|
+
3. Default output path: `<repo>/.metaharness/bench/suite.json` (chosen by upstream).
|
|
31
|
+
4. Suite shape (per upstream): array of `{ input, expectedOutput, weight }` tasks
|
|
32
|
+
derived from existing test cases.
|
|
33
|
+
|
|
34
|
+
### `--op verify`
|
|
35
|
+
1. Resolve `--suite` path; reject if missing.
|
|
36
|
+
2. Shell to `metaharness-darwin bench verify <suite.json>`.
|
|
37
|
+
3. Exit 1 if any task malformed (upstream's signal).
|
|
38
|
+
|
|
39
|
+
## Output shape
|
|
40
|
+
|
|
41
|
+
```json
|
|
42
|
+
{
|
|
43
|
+
"success": true,
|
|
44
|
+
"data": {
|
|
45
|
+
"op": "verify",
|
|
46
|
+
"taskCount": 42,
|
|
47
|
+
"wellFormed": true,
|
|
48
|
+
"durationMs": 870
|
|
49
|
+
}
|
|
50
|
+
}
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
## Exit codes
|
|
54
|
+
|
|
55
|
+
| Code | Meaning |
|
|
56
|
+
|---|---|
|
|
57
|
+
| 0 | OK (or degraded — Darwin absent) |
|
|
58
|
+
| 1 | `--op verify` and suite malformed |
|
|
59
|
+
| 2 | Config error or upstream invocation failure |
|
|
60
|
+
|
|
61
|
+
## Graceful degradation
|
|
62
|
+
|
|
63
|
+
When `@metaharness/darwin` is absent, emits the standard `{degraded: true,
|
|
64
|
+
reason: 'metaharness-darwin-not-available'}` payload and exits 0.
|
|
@@ -1,65 +1,65 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: harness-drift-from-history
|
|
3
|
-
description: One-command drift detection. Composes audit-list + oia-audit + audit-trend into a single primitive — finds the most recent audit in `metaharness-audit` namespace, runs a fresh audit against the current repo, diffs them via ADR-152 §3.1 similarity, and alerts when structural distance crosses `--threshold`. Iter 53 of ADR-150 deep integration.
|
|
4
|
-
argument-hint: "[--path .] [--baseline-since 7d] [--threshold 0.95] [--dry-run] [--format json|table]"
|
|
5
|
-
allowed-tools: Bash
|
|
6
|
-
---
|
|
7
|
-
|
|
8
|
-
The natural ops question after running `oia-audit` weekly is "did anything drift?" Before this skill, the answer required a three-step sequence:
|
|
9
|
-
|
|
10
|
-
```bash
|
|
11
|
-
npx ruflo metaharness audit-list --format json # → pick a key by hand
|
|
12
|
-
npx ruflo metaharness oia-audit --format json > /tmp/curr.json
|
|
13
|
-
npx ruflo metaharness audit-trend \
|
|
14
|
-
--baseline-key <picked-key> --current /tmp/curr.json \
|
|
15
|
-
--alert-on-distance-below 0.95
|
|
16
|
-
```
|
|
17
|
-
|
|
18
|
-
This skill collapses it into one command:
|
|
19
|
-
|
|
20
|
-
```bash
|
|
21
|
-
npx ruflo metaharness drift-from-history --threshold 0.95
|
|
22
|
-
```
|
|
23
|
-
|
|
24
|
-
## What it does
|
|
25
|
-
|
|
26
|
-
1. Lists records from `metaharness-audit` namespace via audit-list.mjs
|
|
27
|
-
2. Picks the most recent record by `startedAt` (or `--baseline-since 7d` skips anything newer than 7 days)
|
|
28
|
-
3. Runs a fresh `oia-audit` against the current path
|
|
29
|
-
4. Diffs the two via audit-trend, applying `--alert-on-distance-below ${threshold}`
|
|
30
|
-
5. Returns the structured drift report
|
|
31
|
-
|
|
32
|
-
## Architectural constraint inheritance (ADR-150)
|
|
33
|
-
|
|
34
|
-
| Constraint | How drift-from-history satisfies it |
|
|
35
|
-
|---|---|
|
|
36
|
-
| Removable | Pure subprocess composition over existing scripts — no new `@metaharness/*` import |
|
|
37
|
-
| Optional | If oia-audit reports `degraded:true`, this skill exits 3 with a degraded payload |
|
|
38
|
-
| Graceful | Empty audit history → exit 2 with hint to seed it; never crashes |
|
|
39
|
-
| CI-gate | Smoke step 17z16 anchors the dispatcher entry + subcommand listing |
|
|
40
|
-
|
|
41
|
-
## Exit codes
|
|
42
|
-
|
|
43
|
-
- 0 — similarity ≥ threshold (or threshold not crossed)
|
|
44
|
-
- 1 — drift detected: similarity < threshold (alert fired)
|
|
45
|
-
- 2 — config error (no history, audit-list failed)
|
|
46
|
-
- 3 — upstream metaharness absent (degraded payload returned)
|
|
47
|
-
|
|
48
|
-
## Example
|
|
49
|
-
|
|
50
|
-
```bash
|
|
51
|
-
$ npx ruflo metaharness drift-from-history --threshold 0.95
|
|
52
|
-
# drift-from-history
|
|
53
|
-
|
|
54
|
-
Baseline: audit-2026-06-16T22-58-47-840Z
|
|
55
|
-
Current: 2026-06-16T23:05:02.231Z
|
|
56
|
-
|
|
57
|
-
Structural similarity: 1 (near-identical)
|
|
58
|
-
Distance: 0
|
|
59
|
-
|
|
60
|
-
✓ similarity ≥ 0.95 — OK
|
|
61
|
-
```
|
|
62
|
-
|
|
63
|
-
## Implementation
|
|
64
|
-
|
|
65
|
-
[`scripts/drift-from-history.mjs`](../../scripts/drift-from-history.mjs)
|
|
1
|
+
---
|
|
2
|
+
name: harness-drift-from-history
|
|
3
|
+
description: One-command drift detection. Composes audit-list + oia-audit + audit-trend into a single primitive — finds the most recent audit in `metaharness-audit` namespace, runs a fresh audit against the current repo, diffs them via ADR-152 §3.1 similarity, and alerts when structural distance crosses `--threshold`. Iter 53 of ADR-150 deep integration.
|
|
4
|
+
argument-hint: "[--path .] [--baseline-since 7d] [--threshold 0.95] [--dry-run] [--format json|table]"
|
|
5
|
+
allowed-tools: Bash
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
The natural ops question after running `oia-audit` weekly is "did anything drift?" Before this skill, the answer required a three-step sequence:
|
|
9
|
+
|
|
10
|
+
```bash
|
|
11
|
+
npx ruflo metaharness audit-list --format json # → pick a key by hand
|
|
12
|
+
npx ruflo metaharness oia-audit --format json > /tmp/curr.json
|
|
13
|
+
npx ruflo metaharness audit-trend \
|
|
14
|
+
--baseline-key <picked-key> --current /tmp/curr.json \
|
|
15
|
+
--alert-on-distance-below 0.95
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
This skill collapses it into one command:
|
|
19
|
+
|
|
20
|
+
```bash
|
|
21
|
+
npx ruflo metaharness drift-from-history --threshold 0.95
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
## What it does
|
|
25
|
+
|
|
26
|
+
1. Lists records from `metaharness-audit` namespace via audit-list.mjs
|
|
27
|
+
2. Picks the most recent record by `startedAt` (or `--baseline-since 7d` skips anything newer than 7 days)
|
|
28
|
+
3. Runs a fresh `oia-audit` against the current path
|
|
29
|
+
4. Diffs the two via audit-trend, applying `--alert-on-distance-below ${threshold}`
|
|
30
|
+
5. Returns the structured drift report
|
|
31
|
+
|
|
32
|
+
## Architectural constraint inheritance (ADR-150)
|
|
33
|
+
|
|
34
|
+
| Constraint | How drift-from-history satisfies it |
|
|
35
|
+
|---|---|
|
|
36
|
+
| Removable | Pure subprocess composition over existing scripts — no new `@metaharness/*` import |
|
|
37
|
+
| Optional | If oia-audit reports `degraded:true`, this skill exits 3 with a degraded payload |
|
|
38
|
+
| Graceful | Empty audit history → exit 2 with hint to seed it; never crashes |
|
|
39
|
+
| CI-gate | Smoke step 17z16 anchors the dispatcher entry + subcommand listing |
|
|
40
|
+
|
|
41
|
+
## Exit codes
|
|
42
|
+
|
|
43
|
+
- 0 — similarity ≥ threshold (or threshold not crossed)
|
|
44
|
+
- 1 — drift detected: similarity < threshold (alert fired)
|
|
45
|
+
- 2 — config error (no history, audit-list failed)
|
|
46
|
+
- 3 — upstream metaharness absent (degraded payload returned)
|
|
47
|
+
|
|
48
|
+
## Example
|
|
49
|
+
|
|
50
|
+
```bash
|
|
51
|
+
$ npx ruflo metaharness drift-from-history --threshold 0.95
|
|
52
|
+
# drift-from-history
|
|
53
|
+
|
|
54
|
+
Baseline: audit-2026-06-16T22-58-47-840Z
|
|
55
|
+
Current: 2026-06-16T23:05:02.231Z
|
|
56
|
+
|
|
57
|
+
Structural similarity: 1 (near-identical)
|
|
58
|
+
Distance: 0
|
|
59
|
+
|
|
60
|
+
✓ similarity ≥ 0.95 — OK
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
## Implementation
|
|
64
|
+
|
|
65
|
+
[`scripts/drift-from-history.mjs`](../../scripts/drift-from-history.mjs)
|
|
@@ -1,131 +1,131 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: harness-evolve
|
|
3
|
-
description: Run `@metaharness/darwin evolve <repo>` to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and promote only measured wins. The model is frozen; the harness evolves. Closes the loop ADR-150 opens (score+genome describe; evolve changes). Degrades gracefully when @metaharness/darwin is absent (ADR-150 + ADR-153 architectural constraints).
|
|
4
|
-
argument-hint: "--repo <path> [--generations 3] [--children 3] [--concurrency 2] [--sandbox real|mock|agent] [--selection pareto|quality-diversity|...] [--mutator deterministic|ruvllm] [--diagnose] [--confirm]"
|
|
5
|
-
allowed-tools: Bash
|
|
6
|
-
---
|
|
7
|
-
|
|
8
|
-
Surfaces the upstream `metaharness-darwin evolve` CLI as a ruflo skill. The
|
|
9
|
-
**write** layer that pairs with ADR-150's read layer (score / genome /
|
|
10
|
-
mcp-scan / threat-model / oia-audit). Use when you have a harness whose
|
|
11
|
-
readiness scores are flat and you want to discover *which* surface mutation
|
|
12
|
-
moves them — without retraining the foundation model.
|
|
13
|
-
|
|
14
|
-
## When to use
|
|
15
|
-
|
|
16
|
-
- A `harness-score` result is below target and you don't know which policy
|
|
17
|
-
surface is responsible.
|
|
18
|
-
- You're seeding a harness for a new vertical and want to find a good
|
|
19
|
-
starting configuration empirically rather than hand-tuning.
|
|
20
|
-
- You're comparing your hand-tuned harness against an evolved baseline
|
|
21
|
-
(treat darwin's champion as the strawman).
|
|
22
|
-
|
|
23
|
-
## When NOT to use
|
|
24
|
-
|
|
25
|
-
- For continuous background optimization. Darwin Mode is human-initiated.
|
|
26
|
-
Wire it into CI for one-shot exploration, not for autonomous self-modification.
|
|
27
|
-
- For ruflo itself in CI. ADR-153 §5 explicitly rejects auto-evolving ruflo
|
|
28
|
-
— the CI gate verifies graceful degradation, not convergence.
|
|
29
|
-
|
|
30
|
-
## Algorithm
|
|
31
|
-
|
|
32
|
-
Implementation: [`scripts/evolve.mjs`](../../scripts/evolve.mjs).
|
|
33
|
-
|
|
34
|
-
1. Validate args (`--repo` exists, caps on `--generations` ≤ 50, `--children`
|
|
35
|
-
≤ 20, `--concurrency` ≤ 8, sandbox/selection/mutator are known values).
|
|
36
|
-
2. Without `--confirm`: print plan + exit 0 (mirrors `harness-mint` safety
|
|
37
|
-
convention; defense in depth over the upstream `safety.ts` checks).
|
|
38
|
-
3. With `--confirm`: shell to `npx -y @metaharness/darwin@~0.8.0 metaharness-darwin evolve <repo> ...`
|
|
39
|
-
via the shared `_darwin.mjs` async helper. Per-generation progress is
|
|
40
|
-
forwarded to stderr; final champion JSON is captured from stdout.
|
|
41
|
-
4. Compute timeout from `generations × children × per-variant` (per-variant
|
|
42
|
-
≈ 60s real, ≈ 2s mock). Caller may override with `--timeout-ms`.
|
|
43
|
-
5. Honor upstream exit code 99 — propagate as "safety-disqualified", do not
|
|
44
|
-
remap. This is a designed-in tripwire (a variant tripped `inspectVariant`
|
|
45
|
-
for secrets / shell-out / network / dynamic-eval). See ADR-153 §"Safety model".
|
|
46
|
-
6. Optional `--alert-on-no-improvement`: exit 1 when champion ≤ parent.
|
|
47
|
-
|
|
48
|
-
## The seven mutation surfaces
|
|
49
|
-
|
|
50
|
-
| Surface | What it owns |
|
|
51
|
-
|---|---|
|
|
52
|
-
| `planner` | task decomposition / step ordering |
|
|
53
|
-
| `contextBuilder` | what gets fed into the prompt |
|
|
54
|
-
| `reviewer` | self-critique / output verification |
|
|
55
|
-
| `retryPolicy` | when + how to retry on failure |
|
|
56
|
-
| `toolPolicy` | which tools the agent may use, under which conditions |
|
|
57
|
-
| `memoryPolicy` | what to persist, recall, forget |
|
|
58
|
-
| `scorePolicy` | how the agent grades its own output |
|
|
59
|
-
|
|
60
|
-
One mutation per variant. Multi-surface mutations are not allowed (causal
|
|
61
|
-
attribution stays clean).
|
|
62
|
-
|
|
63
|
-
## Output
|
|
64
|
-
|
|
65
|
-
Reports land under `<repo>/.metaharness/`:
|
|
66
|
-
|
|
67
|
-
```
|
|
68
|
-
.metaharness/
|
|
69
|
-
archive.json # full lineage tree (sampling next gen draws from this)
|
|
70
|
-
lineage.json # parent→child edges only
|
|
71
|
-
variants/<id>/ # per-variant code (kept for audit)
|
|
72
|
-
runs/<id>/ # per-variant sandbox test output
|
|
73
|
-
reports/winner.json # final champion + score delta vs parent
|
|
74
|
-
```
|
|
75
|
-
|
|
76
|
-
Skill stdout = JSON `{success, data: {champion, plan, durationMs, improved}}`
|
|
77
|
-
(plus `data.diagnosis` when `--diagnose` is passed — see below).
|
|
78
|
-
|
|
79
|
-
## Failure diagnosis (`--diagnose`)
|
|
80
|
-
|
|
81
|
-
GEPA's key trick is natural-language failure diagnosis from execution traces
|
|
82
|
-
feeding the next mutation — not just scalar fitness. `--diagnose` adds a
|
|
83
|
-
modest slice of that: after the evolution completes, the losing / failed
|
|
84
|
-
variants' transcripts are run through darwin's GEPA library ops
|
|
85
|
-
(`analyzeTranscript` + `classifyFailure`, via the shared `importGepa`
|
|
86
|
-
resolver in `scripts/_darwin.mjs`) and a `diagnosis` section is appended to
|
|
87
|
-
the emitted JSON:
|
|
88
|
-
|
|
89
|
-
```json
|
|
90
|
-
"diagnosis": {
|
|
91
|
-
"available": true,
|
|
92
|
-
"scope": "losing-variants",
|
|
93
|
-
"variants": [
|
|
94
|
-
{ "id": "g1_v0", "transcripts": 2,
|
|
95
|
-
"failureClasses": { "exploration-loop": 1, "edit-mechanics": 1 },
|
|
96
|
-
"dominantClass": "exploration-loop" }
|
|
97
|
-
],
|
|
98
|
-
"totals": { "exploration-loop": 1, "edit-mechanics": 1 }
|
|
99
|
-
}
|
|
100
|
-
```
|
|
101
|
-
|
|
102
|
-
Upstream shape caveats (verified against `@metaharness/darwin@0.8.0`):
|
|
103
|
-
|
|
104
|
-
- `metaharness-darwin evolve --json` prints a TEXT leaderboard — the stdout
|
|
105
|
-
carries no JSON and no transcripts. Per-variant run records live at
|
|
106
|
-
`<repo>/.metaharness/runs/<id>.json`.
|
|
107
|
-
- Those run records hold sandbox exec traces (`{taskId, exitCode, stdout,
|
|
108
|
-
stderr}`), which are NOT GEPA `{actionRaw, obs}` transcripts. Diagnosis
|
|
109
|
-
therefore uses GEPA-shaped transcripts when a run record embeds them
|
|
110
|
-
(agent sandbox / future upstream), falls back to the champion's transcript,
|
|
111
|
-
and otherwise emits `diagnosis: {available: false, reason, traceSummary}`
|
|
112
|
-
where `traceSummary` is a mechanical per-variant tally (tasks / failed /
|
|
113
|
-
timedOut / blockedActions).
|
|
114
|
-
- `--diagnose` NEVER fails the run — any internal error degrades to
|
|
115
|
-
`{available: false, reason: "diagnosis-failed: ..."}`.
|
|
116
|
-
|
|
117
|
-
## Exit codes
|
|
118
|
-
|
|
119
|
-
| Code | Meaning |
|
|
120
|
-
|---|---|
|
|
121
|
-
| 0 | Evolved OK, or dry-run, or degraded (Darwin absent) |
|
|
122
|
-
| 1 | `--alert-on-no-improvement` and champion did not beat parent |
|
|
123
|
-
| 2 | Config error or evolution infrastructure failure |
|
|
124
|
-
| 99 | Upstream "safety-disqualified" (PROPAGATED, not remapped) |
|
|
125
|
-
|
|
126
|
-
## Graceful degradation (ADR-150 constraint 3 + ADR-153)
|
|
127
|
-
|
|
128
|
-
When `@metaharness/darwin` is not installed, the script emits
|
|
129
|
-
`{degraded: true, reason: 'metaharness-darwin-not-available', hint: ...}`
|
|
130
|
-
and exits 0. ruflo continues to function. CI's
|
|
131
|
-
`no-metaharness-smoke.yml`-style job asserts this path.
|
|
1
|
+
---
|
|
2
|
+
name: harness-evolve
|
|
3
|
+
description: Run `@metaharness/darwin evolve <repo>` to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and promote only measured wins. The model is frozen; the harness evolves. Closes the loop ADR-150 opens (score+genome describe; evolve changes). Degrades gracefully when @metaharness/darwin is absent (ADR-150 + ADR-153 architectural constraints).
|
|
4
|
+
argument-hint: "--repo <path> [--generations 3] [--children 3] [--concurrency 2] [--sandbox real|mock|agent] [--selection pareto|quality-diversity|...] [--mutator deterministic|ruvllm] [--diagnose] [--confirm]"
|
|
5
|
+
allowed-tools: Bash
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
Surfaces the upstream `metaharness-darwin evolve` CLI as a ruflo skill. The
|
|
9
|
+
**write** layer that pairs with ADR-150's read layer (score / genome /
|
|
10
|
+
mcp-scan / threat-model / oia-audit). Use when you have a harness whose
|
|
11
|
+
readiness scores are flat and you want to discover *which* surface mutation
|
|
12
|
+
moves them — without retraining the foundation model.
|
|
13
|
+
|
|
14
|
+
## When to use
|
|
15
|
+
|
|
16
|
+
- A `harness-score` result is below target and you don't know which policy
|
|
17
|
+
surface is responsible.
|
|
18
|
+
- You're seeding a harness for a new vertical and want to find a good
|
|
19
|
+
starting configuration empirically rather than hand-tuning.
|
|
20
|
+
- You're comparing your hand-tuned harness against an evolved baseline
|
|
21
|
+
(treat darwin's champion as the strawman).
|
|
22
|
+
|
|
23
|
+
## When NOT to use
|
|
24
|
+
|
|
25
|
+
- For continuous background optimization. Darwin Mode is human-initiated.
|
|
26
|
+
Wire it into CI for one-shot exploration, not for autonomous self-modification.
|
|
27
|
+
- For ruflo itself in CI. ADR-153 §5 explicitly rejects auto-evolving ruflo
|
|
28
|
+
— the CI gate verifies graceful degradation, not convergence.
|
|
29
|
+
|
|
30
|
+
## Algorithm
|
|
31
|
+
|
|
32
|
+
Implementation: [`scripts/evolve.mjs`](../../scripts/evolve.mjs).
|
|
33
|
+
|
|
34
|
+
1. Validate args (`--repo` exists, caps on `--generations` ≤ 50, `--children`
|
|
35
|
+
≤ 20, `--concurrency` ≤ 8, sandbox/selection/mutator are known values).
|
|
36
|
+
2. Without `--confirm`: print plan + exit 0 (mirrors `harness-mint` safety
|
|
37
|
+
convention; defense in depth over the upstream `safety.ts` checks).
|
|
38
|
+
3. With `--confirm`: shell to `npx -y @metaharness/darwin@~0.8.0 metaharness-darwin evolve <repo> ...`
|
|
39
|
+
via the shared `_darwin.mjs` async helper. Per-generation progress is
|
|
40
|
+
forwarded to stderr; final champion JSON is captured from stdout.
|
|
41
|
+
4. Compute timeout from `generations × children × per-variant` (per-variant
|
|
42
|
+
≈ 60s real, ≈ 2s mock). Caller may override with `--timeout-ms`.
|
|
43
|
+
5. Honor upstream exit code 99 — propagate as "safety-disqualified", do not
|
|
44
|
+
remap. This is a designed-in tripwire (a variant tripped `inspectVariant`
|
|
45
|
+
for secrets / shell-out / network / dynamic-eval). See ADR-153 §"Safety model".
|
|
46
|
+
6. Optional `--alert-on-no-improvement`: exit 1 when champion ≤ parent.
|
|
47
|
+
|
|
48
|
+
## The seven mutation surfaces
|
|
49
|
+
|
|
50
|
+
| Surface | What it owns |
|
|
51
|
+
|---|---|
|
|
52
|
+
| `planner` | task decomposition / step ordering |
|
|
53
|
+
| `contextBuilder` | what gets fed into the prompt |
|
|
54
|
+
| `reviewer` | self-critique / output verification |
|
|
55
|
+
| `retryPolicy` | when + how to retry on failure |
|
|
56
|
+
| `toolPolicy` | which tools the agent may use, under which conditions |
|
|
57
|
+
| `memoryPolicy` | what to persist, recall, forget |
|
|
58
|
+
| `scorePolicy` | how the agent grades its own output |
|
|
59
|
+
|
|
60
|
+
One mutation per variant. Multi-surface mutations are not allowed (causal
|
|
61
|
+
attribution stays clean).
|
|
62
|
+
|
|
63
|
+
## Output
|
|
64
|
+
|
|
65
|
+
Reports land under `<repo>/.metaharness/`:
|
|
66
|
+
|
|
67
|
+
```
|
|
68
|
+
.metaharness/
|
|
69
|
+
archive.json # full lineage tree (sampling next gen draws from this)
|
|
70
|
+
lineage.json # parent→child edges only
|
|
71
|
+
variants/<id>/ # per-variant code (kept for audit)
|
|
72
|
+
runs/<id>/ # per-variant sandbox test output
|
|
73
|
+
reports/winner.json # final champion + score delta vs parent
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
Skill stdout = JSON `{success, data: {champion, plan, durationMs, improved}}`
|
|
77
|
+
(plus `data.diagnosis` when `--diagnose` is passed — see below).
|
|
78
|
+
|
|
79
|
+
## Failure diagnosis (`--diagnose`)
|
|
80
|
+
|
|
81
|
+
GEPA's key trick is natural-language failure diagnosis from execution traces
|
|
82
|
+
feeding the next mutation — not just scalar fitness. `--diagnose` adds a
|
|
83
|
+
modest slice of that: after the evolution completes, the losing / failed
|
|
84
|
+
variants' transcripts are run through darwin's GEPA library ops
|
|
85
|
+
(`analyzeTranscript` + `classifyFailure`, via the shared `importGepa`
|
|
86
|
+
resolver in `scripts/_darwin.mjs`) and a `diagnosis` section is appended to
|
|
87
|
+
the emitted JSON:
|
|
88
|
+
|
|
89
|
+
```json
|
|
90
|
+
"diagnosis": {
|
|
91
|
+
"available": true,
|
|
92
|
+
"scope": "losing-variants",
|
|
93
|
+
"variants": [
|
|
94
|
+
{ "id": "g1_v0", "transcripts": 2,
|
|
95
|
+
"failureClasses": { "exploration-loop": 1, "edit-mechanics": 1 },
|
|
96
|
+
"dominantClass": "exploration-loop" }
|
|
97
|
+
],
|
|
98
|
+
"totals": { "exploration-loop": 1, "edit-mechanics": 1 }
|
|
99
|
+
}
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
Upstream shape caveats (verified against `@metaharness/darwin@0.8.0`):
|
|
103
|
+
|
|
104
|
+
- `metaharness-darwin evolve --json` prints a TEXT leaderboard — the stdout
|
|
105
|
+
carries no JSON and no transcripts. Per-variant run records live at
|
|
106
|
+
`<repo>/.metaharness/runs/<id>.json`.
|
|
107
|
+
- Those run records hold sandbox exec traces (`{taskId, exitCode, stdout,
|
|
108
|
+
stderr}`), which are NOT GEPA `{actionRaw, obs}` transcripts. Diagnosis
|
|
109
|
+
therefore uses GEPA-shaped transcripts when a run record embeds them
|
|
110
|
+
(agent sandbox / future upstream), falls back to the champion's transcript,
|
|
111
|
+
and otherwise emits `diagnosis: {available: false, reason, traceSummary}`
|
|
112
|
+
where `traceSummary` is a mechanical per-variant tally (tasks / failed /
|
|
113
|
+
timedOut / blockedActions).
|
|
114
|
+
- `--diagnose` NEVER fails the run — any internal error degrades to
|
|
115
|
+
`{available: false, reason: "diagnosis-failed: ..."}`.
|
|
116
|
+
|
|
117
|
+
## Exit codes
|
|
118
|
+
|
|
119
|
+
| Code | Meaning |
|
|
120
|
+
|---|---|
|
|
121
|
+
| 0 | Evolved OK, or dry-run, or degraded (Darwin absent) |
|
|
122
|
+
| 1 | `--alert-on-no-improvement` and champion did not beat parent |
|
|
123
|
+
| 2 | Config error or evolution infrastructure failure |
|
|
124
|
+
| 99 | Upstream "safety-disqualified" (PROPAGATED, not remapped) |
|
|
125
|
+
|
|
126
|
+
## Graceful degradation (ADR-150 constraint 3 + ADR-153)
|
|
127
|
+
|
|
128
|
+
When `@metaharness/darwin` is not installed, the script emits
|
|
129
|
+
`{degraded: true, reason: 'metaharness-darwin-not-available', hint: ...}`
|
|
130
|
+
and exits 0. ruflo continues to function. CI's
|
|
131
|
+
`no-metaharness-smoke.yml`-style job asserts this path.
|
|
@@ -1,54 +1,54 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: harness-genome
|
|
3
|
-
description: 7-section repo readiness report from `metaharness genome <path>`. Returns repo_type / agent_topology / risk_score / mcp_surface / test_confidence / publish_readiness. Pure-read; degrades gracefully (ADR-150).
|
|
4
|
-
argument-hint: "[--path .] [--alert-on-risk-above 0.5] [--format table|json]"
|
|
5
|
-
allowed-tools: Bash
|
|
6
|
-
---
|
|
7
|
-
|
|
8
|
-
Companion to `harness-score`. Where score is a 5-dimension numeric
|
|
9
|
-
scorecard, genome is a 7-section categorical/numeric report covering
|
|
10
|
-
repo type, agent topology recommendations, risk score (0-1), MCP
|
|
11
|
-
surface area, test confidence (0-1), and publish readiness (0-1).
|
|
12
|
-
|
|
13
|
-
## Algorithm
|
|
14
|
-
|
|
15
|
-
Implementation: [`scripts/genome.mjs`](../../scripts/genome.mjs).
|
|
16
|
-
|
|
17
|
-
1. Shell out to `npx metaharness genome <path> --json` (60s hard timeout).
|
|
18
|
-
2. Parse the shape: `{ repo_type, agent_topology[], risk_score,
|
|
19
|
-
mcp_surface, test_confidence, publish_readiness }`.
|
|
20
|
-
3. If `--alert-on-risk-above N`: exit 1 when `risk_score > N`.
|
|
21
|
-
4. Output JSON (default) or markdown.
|
|
22
|
-
|
|
23
|
-
## Phase-0 baseline (ruflo, measured 2026-06-16)
|
|
24
|
-
|
|
25
|
-
```
|
|
26
|
-
{
|
|
27
|
-
"repo_type": "node_mcp_ci",
|
|
28
|
-
"agent_topology": ["maintainer", "tester", "security", "release"],
|
|
29
|
-
"risk_score": 0.27,
|
|
30
|
-
"mcp_surface": "remote",
|
|
31
|
-
"test_confidence": 0.8,
|
|
32
|
-
"publish_readiness": 0.9
|
|
33
|
-
}
|
|
34
|
-
```
|
|
35
|
-
|
|
36
|
-
Ruflo's `risk_score: 0.27` is low (good). `publish_readiness: 0.9` is
|
|
37
|
-
high. The `mcp_surface: "remote"` reflects that ruflo's MCP servers are
|
|
38
|
-
hosted, not bundled.
|
|
39
|
-
|
|
40
|
-
## When to use
|
|
41
|
-
|
|
42
|
-
- Pre-mint review: "before scaffolding a custom harness from this repo,
|
|
43
|
-
should we?" — genome answers it categorically.
|
|
44
|
-
- Drift detection: capture genome snapshots over time, diff via
|
|
45
|
-
cost-diff-style tooling to spot when `agent_topology` recommendations
|
|
46
|
-
drift away from a deliberate architecture choice.
|
|
47
|
-
- CI gate: `--alert-on-risk-above 0.5` fails the build when the repo's
|
|
48
|
-
risk profile crosses a threshold.
|
|
49
|
-
|
|
50
|
-
## Pairs with
|
|
51
|
-
|
|
52
|
-
- `harness-score` — numeric readiness
|
|
53
|
-
- `harness-mcp-scan` — static MCP security findings
|
|
54
|
-
- `harness-threat-model` — enterprise-review-grade threat model
|
|
1
|
+
---
|
|
2
|
+
name: harness-genome
|
|
3
|
+
description: 7-section repo readiness report from `metaharness genome <path>`. Returns repo_type / agent_topology / risk_score / mcp_surface / test_confidence / publish_readiness. Pure-read; degrades gracefully (ADR-150).
|
|
4
|
+
argument-hint: "[--path .] [--alert-on-risk-above 0.5] [--format table|json]"
|
|
5
|
+
allowed-tools: Bash
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
Companion to `harness-score`. Where score is a 5-dimension numeric
|
|
9
|
+
scorecard, genome is a 7-section categorical/numeric report covering
|
|
10
|
+
repo type, agent topology recommendations, risk score (0-1), MCP
|
|
11
|
+
surface area, test confidence (0-1), and publish readiness (0-1).
|
|
12
|
+
|
|
13
|
+
## Algorithm
|
|
14
|
+
|
|
15
|
+
Implementation: [`scripts/genome.mjs`](../../scripts/genome.mjs).
|
|
16
|
+
|
|
17
|
+
1. Shell out to `npx metaharness genome <path> --json` (60s hard timeout).
|
|
18
|
+
2. Parse the shape: `{ repo_type, agent_topology[], risk_score,
|
|
19
|
+
mcp_surface, test_confidence, publish_readiness }`.
|
|
20
|
+
3. If `--alert-on-risk-above N`: exit 1 when `risk_score > N`.
|
|
21
|
+
4. Output JSON (default) or markdown.
|
|
22
|
+
|
|
23
|
+
## Phase-0 baseline (ruflo, measured 2026-06-16)
|
|
24
|
+
|
|
25
|
+
```
|
|
26
|
+
{
|
|
27
|
+
"repo_type": "node_mcp_ci",
|
|
28
|
+
"agent_topology": ["maintainer", "tester", "security", "release"],
|
|
29
|
+
"risk_score": 0.27,
|
|
30
|
+
"mcp_surface": "remote",
|
|
31
|
+
"test_confidence": 0.8,
|
|
32
|
+
"publish_readiness": 0.9
|
|
33
|
+
}
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
Ruflo's `risk_score: 0.27` is low (good). `publish_readiness: 0.9` is
|
|
37
|
+
high. The `mcp_surface: "remote"` reflects that ruflo's MCP servers are
|
|
38
|
+
hosted, not bundled.
|
|
39
|
+
|
|
40
|
+
## When to use
|
|
41
|
+
|
|
42
|
+
- Pre-mint review: "before scaffolding a custom harness from this repo,
|
|
43
|
+
should we?" — genome answers it categorically.
|
|
44
|
+
- Drift detection: capture genome snapshots over time, diff via
|
|
45
|
+
cost-diff-style tooling to spot when `agent_topology` recommendations
|
|
46
|
+
drift away from a deliberate architecture choice.
|
|
47
|
+
- CI gate: `--alert-on-risk-above 0.5` fails the build when the repo's
|
|
48
|
+
risk profile crosses a threshold.
|
|
49
|
+
|
|
50
|
+
## Pairs with
|
|
51
|
+
|
|
52
|
+
- `harness-score` — numeric readiness
|
|
53
|
+
- `harness-mcp-scan` — static MCP security findings
|
|
54
|
+
- `harness-threat-model` — enterprise-review-grade threat model
|