session-orchestrator 5.1.0 → 5.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/architecture/SKILL.md +3 -1
- package/.agents/skills/autopilot/SKILL.md +6 -1
- package/.agents/skills/autopilot/agents/openai.yaml +5 -0
- package/.agents/skills/bootstrap/SKILL.md +7 -1
- package/.agents/skills/bootstrap/agents/openai.yaml +5 -0
- package/.agents/skills/brainstorm/SKILL.md +8 -1
- package/.agents/skills/brainstorm/agents/openai.yaml +5 -0
- package/.agents/skills/claude-md-drift-check/SKILL.md +3 -1
- package/.agents/skills/close/SKILL.md +21 -0
- package/.agents/skills/close/agents/openai.yaml +5 -0
- package/.agents/skills/convergence-monitoring/SKILL.md +4 -2
- package/.agents/skills/debug/SKILL.md +7 -1
- package/.agents/skills/debug/agents/openai.yaml +5 -0
- package/.agents/skills/discovery/SKILL.md +7 -2
- package/.agents/skills/discovery/agents/openai.yaml +5 -0
- package/.agents/skills/dispatcher/SKILL.md +7 -1
- package/.agents/skills/dispatcher/agents/openai.yaml +5 -0
- package/.agents/skills/docs-orchestrator/SKILL.md +3 -1
- package/.agents/skills/ecosystem-health/SKILL.md +3 -1
- package/.agents/skills/eli5/SKILL.md +7 -1
- package/.agents/skills/eli5/agents/openai.yaml +5 -0
- package/.agents/skills/eval/SKILL.md +7 -2
- package/.agents/skills/eval/agents/openai.yaml +5 -0
- package/.agents/skills/evolve/SKILL.md +8 -3
- package/.agents/skills/evolve/agents/openai.yaml +5 -0
- package/.agents/skills/frontmatter-guard/SKILL.md +3 -1
- package/.agents/skills/gitlab-ops/SKILL.md +3 -1
- package/.agents/skills/gitlab-portfolio/SKILL.md +3 -1
- package/.agents/skills/go/SKILL.md +22 -0
- package/.agents/skills/go/agents/openai.yaml +5 -0
- package/.agents/skills/grill/SKILL.md +7 -1
- package/.agents/skills/grill/agents/openai.yaml +5 -0
- package/.agents/skills/harness-audit/SKILL.md +20 -0
- package/.agents/skills/harness-audit/agents/openai.yaml +5 -0
- package/.agents/skills/hook-development/SKILL.md +3 -1
- package/.agents/skills/mcp-builder/SKILL.md +3 -1
- package/.agents/skills/memory-cleanup/SKILL.md +6 -1
- package/.agents/skills/memory-cleanup/agents/openai.yaml +5 -0
- package/.agents/skills/mode-selector/SKILL.md +3 -1
- package/.agents/skills/npm-publish/SKILL.md +4 -2
- package/.agents/skills/peekaboo-driver/SKILL.md +3 -1
- package/.agents/skills/persona-panel/SKILL.md +6 -1
- package/.agents/skills/persona-panel/agents/openai.yaml +5 -0
- package/.agents/skills/plan/SKILL.md +8 -2
- package/.agents/skills/plan/agents/openai.yaml +5 -0
- package/.agents/skills/playwright-driver/SKILL.md +3 -1
- package/.agents/skills/portfolio/SKILL.md +21 -0
- package/.agents/skills/portfolio/agents/openai.yaml +5 -0
- package/.agents/skills/quality-gates/SKILL.md +3 -1
- package/.agents/skills/reconcile/SKILL.md +6 -1
- package/.agents/skills/reconcile/agents/openai.yaml +5 -0
- package/.agents/skills/release/SKILL.md +22 -0
- package/.agents/skills/release/agents/openai.yaml +5 -0
- package/.agents/skills/remote-offload/SKILL.md +3 -1
- package/.agents/skills/repo-audit/SKILL.md +6 -1
- package/.agents/skills/repo-audit/agents/openai.yaml +5 -0
- package/.agents/skills/session/SKILL.md +21 -0
- package/.agents/skills/session/agents/openai.yaml +5 -0
- package/.agents/skills/session-end/SKILL.md +3 -1
- package/.agents/skills/session-plan/SKILL.md +3 -1
- package/.agents/skills/session-start/SKILL.md +3 -1
- package/.agents/skills/spinout/SKILL.md +6 -1
- package/.agents/skills/spinout/agents/openai.yaml +5 -0
- package/.agents/skills/sunset-review/SKILL.md +7 -1
- package/.agents/skills/sunset-review/agents/openai.yaml +5 -0
- package/.agents/skills/templates-ack/SKILL.md +21 -0
- package/.agents/skills/templates-ack/agents/openai.yaml +5 -0
- package/.agents/skills/test/SKILL.md +21 -0
- package/.agents/skills/test/agents/openai.yaml +5 -0
- package/.agents/skills/test-runner/SKILL.md +3 -1
- package/.agents/skills/tmux-layout/SKILL.md +3 -1
- package/.agents/skills/using-orchestrator/SKILL.md +3 -1
- package/.agents/skills/ux-grill/SKILL.md +7 -1
- package/.agents/skills/ux-grill/agents/openai.yaml +5 -0
- package/.agents/skills/vault-mirror/SKILL.md +3 -1
- package/.agents/skills/vault-sync/SKILL.md +3 -1
- package/.agents/skills/wave-executor/SKILL.md +3 -1
- package/.agents/skills/write-executable-plan/SKILL.md +3 -1
- package/.claude-plugin/marketplace.json +1 -1
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +4 -4
- package/.codex-plugin/skills/autopilot/SKILL.md +5 -4
- package/.codex-plugin/skills/bootstrap/SKILL.md +8 -4
- package/.codex-plugin/skills/brainstorm/SKILL.md +11 -4
- package/.codex-plugin/skills/close/SKILL.md +3 -3
- package/.codex-plugin/skills/convergence-monitoring/SKILL.md +1 -1
- package/.codex-plugin/skills/debug/SKILL.md +11 -4
- package/.codex-plugin/skills/discovery/SKILL.md +8 -4
- package/.codex-plugin/skills/dispatcher/SKILL.md +4 -4
- package/.codex-plugin/skills/eli5/SKILL.md +9 -4
- package/.codex-plugin/skills/eval/SKILL.md +9 -4
- package/.codex-plugin/skills/evolve/SKILL.md +9 -4
- package/.codex-plugin/skills/go/SKILL.md +3 -3
- package/.codex-plugin/skills/grill/SKILL.md +11 -4
- package/.codex-plugin/skills/harness-audit/SKILL.md +4 -3
- package/.codex-plugin/skills/memory-cleanup/SKILL.md +9 -4
- package/.codex-plugin/skills/npm-publish/SKILL.md +1 -1
- package/.codex-plugin/skills/persona-panel/SKILL.md +5 -5
- package/.codex-plugin/skills/plan/SKILL.md +8 -4
- package/.codex-plugin/skills/portfolio/SKILL.md +3 -3
- package/.codex-plugin/skills/reconcile/SKILL.md +9 -4
- package/.codex-plugin/skills/release/SKILL.md +3 -3
- package/.codex-plugin/skills/repo-audit/SKILL.md +6 -4
- package/.codex-plugin/skills/session/SKILL.md +1 -1
- package/.codex-plugin/skills/spinout/SKILL.md +4 -4
- package/.codex-plugin/skills/sunset-review/SKILL.md +5 -4
- package/.codex-plugin/skills/test/SKILL.md +3 -3
- package/.codex-plugin/skills/ux-grill/SKILL.md +11 -4
- package/.cursor/commands/autopilot.md +4 -4
- package/.cursor/commands/bootstrap.md +5 -4
- package/.cursor/commands/brainstorm.md +5 -4
- package/.cursor/commands/close.md +4 -3
- package/.cursor/commands/debug.md +4 -4
- package/.cursor/commands/discovery.md +4 -4
- package/.cursor/commands/dispatcher.md +4 -4
- package/.cursor/commands/eli5.md +4 -4
- package/.cursor/commands/eval.md +4 -4
- package/.cursor/commands/evolve.md +4 -4
- package/.cursor/commands/go.md +4 -3
- package/.cursor/commands/grill.md +4 -4
- package/.cursor/commands/harness-audit.md +3 -3
- package/.cursor/commands/memory-cleanup.md +4 -4
- package/.cursor/commands/persona-panel.md +4 -4
- package/.cursor/commands/plan.md +5 -4
- package/.cursor/commands/portfolio.md +3 -3
- package/.cursor/commands/reconcile.md +4 -4
- package/.cursor/commands/release.md +4 -3
- package/.cursor/commands/repo-audit.md +4 -4
- package/.cursor/commands/session.md +1 -1
- package/.cursor/commands/spinout.md +4 -4
- package/.cursor/commands/sunset-review.md +4 -4
- package/.cursor/commands/test.md +3 -3
- package/.cursor/commands/ux-grill.md +4 -4
- package/.cursor/rules/000-session-orchestrator.mdc +0 -2
- package/.cursor/rules/010-session-workflow.mdc +2 -2
- package/.cursor/rules/050-plan.mdc +1 -1
- package/.cursor/skills/bootstrap/SKILL.md +1 -0
- package/.cursor/skills/close/SKILL.md +13 -0
- package/.cursor/skills/convergence-monitoring/SKILL.md +1 -0
- package/.cursor/skills/debug/SKILL.md +0 -1
- package/.cursor/skills/discovery/SKILL.md +0 -1
- package/.cursor/skills/dispatcher/SKILL.md +0 -1
- package/.cursor/skills/eli5/SKILL.md +0 -1
- package/.cursor/skills/eval/SKILL.md +1 -1
- package/.cursor/skills/evolve/SKILL.md +0 -1
- package/.cursor/skills/go/SKILL.md +13 -0
- package/.cursor/skills/grill/SKILL.md +0 -1
- package/.cursor/skills/harness-audit/SKILL.md +12 -0
- package/.cursor/skills/npm-publish/SKILL.md +1 -0
- package/.cursor/skills/portfolio/SKILL.md +12 -0
- package/.cursor/skills/release/SKILL.md +13 -0
- package/.cursor/skills/repo-audit/SKILL.md +0 -1
- package/.cursor/skills/sunset-review/SKILL.md +0 -1
- package/.cursor/skills/test/SKILL.md +12 -0
- package/.cursor/skills/ux-grill/SKILL.md +0 -1
- package/.cursor-plugin/plugin.json +1 -1
- package/.orchestrator/policy/blocked-commands.json +13 -4
- package/AGENTS.md +3 -2
- package/CHANGELOG.md +197 -0
- package/README.md +11 -9
- package/SECURITY.md +12 -0
- package/agents/dialectic-deriver.md +13 -10
- package/agents/eval-judge.md +67 -45
- package/agents/skill-applied-judge.md +34 -19
- package/commands/session.md +17 -3
- package/docs/baseline.md +12 -6
- package/docs/ci-setup.md +53 -0
- package/docs/codex-setup.md +15 -3
- package/docs/components.md +13 -6
- package/docs/events-schema.md +59 -9
- package/docs/install.md +16 -0
- package/docs/persona-panel.md +1 -1
- package/docs/pi-setup.md +1 -1
- package/docs/rule-authoring.md +135 -14
- package/docs/scope-collision-guard.md +2 -0
- package/docs/session-config-reference.md +106 -11
- package/docs/session-config-template.md +31 -2
- package/docs/telemetry.md +2 -0
- package/hooks/_lib/hook-import-set.json +125 -8
- package/hooks/_lib/subagent-paths.mjs +15 -0
- package/hooks/_lib/subagent-transcript.mjs +582 -31
- package/hooks/_lib/vcs-create-matcher.mjs +217 -62
- package/hooks/config-protection.mjs +11 -3
- package/hooks/cwd-change-restore.mjs +11 -3
- package/hooks/enforce-commands.mjs +70 -23
- package/hooks/enforce-scope.mjs +143 -33
- package/hooks/hooks-codex.json +1 -1
- package/hooks/hooks.json +1 -1
- package/hooks/loop-guard.mjs +11 -3
- package/hooks/on-session-end.mjs +72 -25
- package/hooks/on-session-start.mjs +48 -11
- package/hooks/on-stop.mjs +211 -23
- package/hooks/operator-steer.mjs +11 -3
- package/hooks/post-bash-issue-budget-refund.mjs +18 -8
- package/hooks/post-bash-write-verify.mjs +6 -2
- package/hooks/post-edit-import-probe.mjs +17 -9
- package/hooks/post-edit-validate.mjs +13 -5
- package/hooks/post-subagent-discovery-validator.mjs +98 -13
- package/hooks/post-tool-batch-wave-signal.mjs +200 -38
- package/hooks/post-tool-failure-corrective-context.mjs +11 -5
- package/hooks/post-tooluse-frontend-slop.mjs +10 -4
- package/hooks/pre-auq-clarity.mjs +18 -2
- package/hooks/pre-bash-destructive-guard.mjs +80 -9
- package/hooks/pre-bash-issue-budget.mjs +119 -28
- package/hooks/pre-bash-memory-propose-audit.mjs +86 -54
- package/hooks/pre-bash-sessions-ledger-guard.mjs +391 -20
- package/hooks/pre-bash-staging-fence.mjs +335 -31
- package/hooks/pre-bash-templates-first.mjs +19 -14
- package/hooks/pre-task-scope-disjoint.mjs +385 -5
- package/hooks/skill-invocation-telemetry.mjs +2 -1
- package/hooks/subagent-telemetry.mjs +15 -19
- package/hooks/wave-scope-commit-guard.mjs +197 -100
- package/monitors/monitors.json +1 -1
- package/output-styles/wave-summary.md +1 -1
- package/package.json +2 -1
- package/pi/prompts/autopilot.md +3 -3
- package/pi/prompts/bootstrap.md +3 -3
- package/pi/prompts/brainstorm.md +3 -3
- package/pi/prompts/close.md +2 -2
- package/pi/prompts/debug.md +3 -3
- package/pi/prompts/discovery.md +3 -3
- package/pi/prompts/dispatcher.md +3 -3
- package/pi/prompts/eli5.md +3 -3
- package/pi/prompts/eval.md +3 -3
- package/pi/prompts/evolve.md +3 -3
- package/pi/prompts/go.md +2 -2
- package/pi/prompts/grill.md +3 -3
- package/pi/prompts/harness-audit.md +2 -3
- package/pi/prompts/memory-cleanup.md +3 -3
- package/pi/prompts/persona-panel.md +3 -3
- package/pi/prompts/plan.md +3 -3
- package/pi/prompts/portfolio.md +2 -2
- package/pi/prompts/reconcile.md +3 -3
- package/pi/prompts/release.md +3 -3
- package/pi/prompts/repo-audit.md +3 -4
- package/pi/prompts/session.md +2 -2
- package/pi/prompts/spinout.md +3 -3
- package/pi/prompts/sunset-review.md +3 -3
- package/pi/prompts/templates-ack.md +1 -1
- package/pi/prompts/test.md +3 -3
- package/pi/prompts/ux-grill.md +3 -3
- package/rules/README.md +1 -1
- package/rules/opt-in-domain/prompt-caching.md +1 -1
- package/rules/opt-in-stack/backend-data.md +1 -1
- package/rules/opt-in-stack/backend.md +3 -3
- package/rules/opt-in-stack/frontend.md +1 -1
- package/rules/opt-in-stack/security-web.md +3 -3
- package/rules/opt-in-stack/swift.md +1 -1
- package/scripts/archive-closed-prds.mjs +2 -2
- package/scripts/auq-audit.mjs +2 -3
- package/scripts/autopilot.mjs +23 -2
- package/scripts/backfill-abandoned-sessions.mjs +171 -15
- package/scripts/backfill-evidence-digest.mjs +2 -1
- package/scripts/backfill-learnings-from-vault.mjs +2 -2
- package/scripts/check-package-manager.mjs +2 -2
- package/scripts/check-sessions-integrity.mjs +300 -0
- package/scripts/ci/assert-vitest-green.mjs +2 -1
- package/scripts/dialectic-deriver.mjs +50 -13
- package/scripts/emit-session.mjs +77 -32
- package/scripts/eval-session.mjs +65 -3
- package/scripts/export-hw-learnings.mjs +2 -1
- package/scripts/express-path.mjs +1 -1
- package/scripts/gc-stale-worktrees.mjs +2 -1
- package/scripts/generate-agents-skills.mjs +102 -29
- package/scripts/generate-codex-skills.mjs +48 -4
- package/scripts/generate-cursor-adapter.mjs +220 -11
- package/scripts/generate-hook-import-set.mjs +12 -27
- package/scripts/generate-pi-prompts.mjs +183 -13
- package/scripts/github-protection-audit.mjs +2 -3
- package/scripts/lib/agent-frontmatter.mjs +23 -1
- package/scripts/lib/agent-status.mjs +2 -31
- package/scripts/lib/auq/clarity.mjs +10 -2
- package/scripts/lib/auq/parse.mjs +12 -31
- package/scripts/lib/auq/schema.mjs +56 -41
- package/scripts/lib/auto-dialectic.mjs +304 -15
- package/scripts/lib/autopilot/flags.mjs +12 -1
- package/scripts/lib/autopilot/kill-switches.mjs +6 -3
- package/scripts/lib/autopilot/loop.mjs +14 -1
- package/scripts/lib/autopilot/stall-sampler.mjs +80 -23
- package/scripts/lib/ci-status-banner.mjs +376 -16
- package/scripts/lib/claude-md-budget-lint.mjs +2 -5
- package/scripts/lib/command-blocker.mjs +408 -33
- package/scripts/lib/config/dialectic.mjs +12 -3
- package/scripts/lib/config/drift-check.mjs +19 -0
- package/scripts/lib/config/gate.mjs +74 -0
- package/scripts/lib/config/reaper.mjs +162 -0
- package/scripts/lib/config.mjs +14 -0
- package/scripts/lib/convergence-monitor.mjs +76 -13
- package/scripts/lib/cursor-hook-bridge.mjs +2 -2
- package/scripts/lib/description-surface.mjs +2 -5
- package/scripts/lib/dispatcher/cli.mjs +2 -1
- package/scripts/lib/ecosystem-health.mjs +11 -0
- package/scripts/lib/ecosystem-wizard.mjs +2 -1
- package/scripts/lib/eval/engine.mjs +421 -53
- package/scripts/lib/eval/judge.mjs +463 -40
- package/scripts/lib/eval/schema.mjs +10 -1
- package/scripts/lib/events-rotation.mjs +221 -25
- package/scripts/lib/events-schema.mjs +114 -0
- package/scripts/lib/events.mjs +524 -5
- package/scripts/lib/fetch-baseline.mjs +3 -8
- package/scripts/lib/frontmatter-guard.mjs +21 -10
- package/scripts/lib/gates/gate-baseline.mjs +27 -2
- package/scripts/lib/gates/gate-full.mjs +28 -3
- package/scripts/lib/gates/gate-helpers.mjs +243 -21
- package/scripts/lib/gates/gate-incremental.mjs +28 -3
- package/scripts/lib/gates/gate-per-file.mjs +27 -2
- package/scripts/lib/gitlab-ops/stale-mr-sweep.mjs +2 -1
- package/scripts/lib/gitlab-portfolio/cli.mjs +2 -1
- package/scripts/lib/gitlab-portfolio/markdown-writer.mjs +6 -1
- package/scripts/lib/instruction-budget-guard.mjs +332 -50
- package/scripts/lib/io.mjs +42 -8
- package/scripts/lib/is-main-module.mjs +82 -0
- package/scripts/lib/issue-close-strip-labels.mjs +207 -49
- package/scripts/lib/js-mask.mjs +197 -0
- package/scripts/lib/learnings/evolve-telemetry.mjs +11 -7
- package/scripts/lib/locks/index.mjs +32 -25
- package/scripts/lib/maintenance-due-banner.mjs +122 -91
- package/scripts/lib/orphan-reaper.mjs +1588 -0
- package/scripts/lib/peer-cards/merger.mjs +48 -10
- package/scripts/lib/peer-cards/reader.mjs +78 -2
- package/scripts/lib/peer-discovery.mjs +2 -5
- package/scripts/lib/playwright-driver/runner.mjs +2 -1
- package/scripts/lib/process-group.mjs +899 -0
- package/scripts/lib/quality-gate.mjs +107 -28
- package/scripts/lib/reconcile/backlog.mjs +368 -0
- package/scripts/lib/reconcile/engine.mjs +55 -188
- package/scripts/lib/reconcile/rule-expiry-sweep.mjs +884 -0
- package/scripts/lib/reconcile/sanitize.mjs +69 -3
- package/scripts/lib/reconcile-nudge-banner.mjs +138 -45
- package/scripts/lib/resource-probe/parsers.mjs +31 -0
- package/scripts/lib/rule-loader.mjs +41 -12
- package/scripts/lib/rules-sync.mjs +2 -5
- package/scripts/lib/scope-echo.mjs +429 -7
- package/scripts/lib/scope-gate.mjs +605 -1
- package/scripts/lib/session-close-backfill.mjs +91 -12
- package/scripts/lib/session-id.mjs +9 -20
- package/scripts/lib/session-invocation.mjs +20 -0
- package/scripts/lib/session-schema/constants.mjs +30 -2
- package/scripts/lib/session-schema/normalizer.mjs +56 -4
- package/scripts/lib/session-schema.mjs +8 -3
- package/scripts/lib/session-start-probes.mjs +95 -10
- package/scripts/lib/sessions-canonical.mjs +23 -0
- package/scripts/lib/sessions-integrity-banner.mjs +7 -1
- package/scripts/lib/sessions-staleness-banner.mjs +193 -51
- package/scripts/lib/skill-evidence-window.mjs +891 -0
- package/scripts/lib/skill-evolution/candidate-intake.mjs +133 -12
- package/scripts/lib/skill-evolution/engine.mjs +18 -9
- package/scripts/lib/skill-judge.mjs +45 -3
- package/scripts/lib/state-md.mjs +84 -3
- package/scripts/lib/sunset/walker.mjs +31 -4
- package/scripts/lib/tail-window.mjs +56 -0
- package/scripts/lib/telemetry/schema.mjs +30 -0
- package/scripts/lib/telemetry/sync.mjs +61 -6
- package/scripts/lib/telemetry-flush-health-banner.mjs +4 -22
- package/scripts/lib/test-runner/issue-reconcile.mjs +48 -16
- package/scripts/lib/tests-src-ratio.mjs +2 -6
- package/scripts/lib/tmux-layout/telemetry-stats.mjs +74 -14
- package/scripts/lib/user-invocable-skills.mjs +205 -0
- package/scripts/lib/ux-grill/reconcile.mjs +48 -22
- package/scripts/lib/validate/check-agents-skills.mjs +26 -15
- package/scripts/lib/validate/check-banner-parity.mjs +2 -2
- package/scripts/lib/validate/check-cursor-adapter.mjs +3 -2
- package/scripts/lib/validate/check-dead-bridge.mjs +2 -2
- package/scripts/lib/validate/check-doc-cli-commands.mjs +2 -2
- package/scripts/lib/validate/check-entry-guard.mjs +329 -0
- package/scripts/lib/validate/check-guard-requires-parity.mjs +2 -2
- package/scripts/lib/validate/check-hook-entry-guards.mjs +636 -0
- package/scripts/lib/validate/check-hooks-emit-event-guard.mjs +2 -2
- package/scripts/lib/validate/check-learning-provenance.mjs +2 -2
- package/scripts/lib/validate/check-pi-prompts.mjs +1 -0
- package/scripts/lib/validate/check-rules.mjs +7 -5
- package/scripts/lib/validate/check-skill-links.mjs +35 -6
- package/scripts/lib/validate/check-skill-script-paths.mjs +241 -29
- package/scripts/lib/validate/check-test-git-config-target.mjs +26 -36
- package/scripts/lib/validate/check-unicode-safety.mjs +2 -2
- package/scripts/lib/validate/check-untracked-test-deps.mjs +9 -104
- package/scripts/lib/validate/check-unwired-features.mjs +220 -33
- package/scripts/lib/validate/check-validator-registration.mjs +36 -12
- package/scripts/lib/validate/check-vcs-repo-flag.mjs +2 -2
- package/scripts/lib/validate/confidential-names.mjs +10 -0
- package/scripts/lib/validate-vendored-rules.mjs +39 -12
- package/scripts/lib/vault-mirror/namespace.mjs +46 -8
- package/scripts/lib/vault-mirror/process.mjs +10 -3
- package/scripts/lib/vault-mirror/render-sessions.mjs +12 -2
- package/scripts/lib/vault-status/narrative-mirror.mjs +31 -7
- package/scripts/lib/vault-yaml.mjs +118 -0
- package/scripts/lib/wave-transcript-tail.mjs +2 -2
- package/scripts/lib/worktree/lifecycle.mjs +153 -1
- package/scripts/lock-reaper.mjs +2 -1
- package/scripts/materialize-wave-scope.mjs +87 -4
- package/scripts/migrate-sessions-jsonl.mjs +2 -1
- package/scripts/migrate-vault-paths.mjs +2 -3
- package/scripts/release-session-lock.mjs +305 -0
- package/scripts/release.mjs +109 -39
- package/scripts/relocate-vault-corpus.mjs +2 -3
- package/scripts/repair-invalid-sessions.mjs +2 -2
- package/scripts/resolve-session-invocation.mjs +59 -0
- package/scripts/run-quality-gate.mjs +156 -17
- package/scripts/session-shape.mjs +2 -2
- package/scripts/site-numbers.mjs +35 -11
- package/scripts/sweep-expired-rules.mjs +227 -0
- package/scripts/validate-plugin.mjs +21 -0
- package/scripts/validate-wave-scope.mjs +32 -105
- package/scripts/vault-consolidate.mjs +2 -2
- package/scripts/vault-mirror.mjs +11 -4
- package/scripts/wave-scope-binding.mjs +2 -3
- package/skills/_shared/bootstrap-gate.md +1 -1
- package/skills/_shared/monitor-patterns.md +1 -1
- package/skills/_shared/platform-tools.md +23 -11
- package/skills/_shared/research-evidence.md +53 -0
- package/skills/_shared/state-ownership.md +3 -0
- package/skills/autopilot/SKILL.md +80 -11
- package/skills/bootstrap/SKILL.md +51 -1
- package/skills/brainstorm/SKILL.md +16 -0
- package/skills/claude-md-drift-check/SKILL.md +1 -1
- package/skills/claude-md-drift-check/checker.mjs +49 -11
- package/{commands/close.md → skills/close/SKILL.md} +9 -3
- package/skills/convergence-monitoring/README.md +8 -1
- package/skills/convergence-monitoring/SIGNALS.md +50 -6
- package/skills/convergence-monitoring/SKILL.md +15 -6
- package/skills/debug/SKILL.md +10 -0
- package/skills/discovery/SKILL.md +24 -1
- package/skills/discovery/probes-session.md +2 -2
- package/skills/dispatcher/SKILL.md +38 -7
- package/skills/eli5/SKILL.md +11 -0
- package/skills/eval/SKILL.md +52 -23
- package/skills/eval/rubric-v1.md +1 -0
- package/skills/eval/rubric-v2.md +457 -0
- package/skills/evolve/SKILL.md +9 -2
- package/skills/evolve/references/evolve-dialectic-mode.md +46 -25
- package/skills/gitlab-ops/SKILL.md +3 -2
- package/{commands/go.md → skills/go/SKILL.md} +9 -1
- package/skills/grill/SKILL.md +19 -0
- package/{commands/harness-audit.md → skills/harness-audit/SKILL.md} +7 -2
- package/skills/hook-development/SKILL.md +46 -41
- package/skills/memory-cleanup/SKILL.md +7 -0
- package/skills/npm-publish/SKILL.md +2 -2
- package/skills/persona-panel/SKILL.md +56 -1
- package/skills/persona-panel/persona-format.md +1 -1
- package/skills/plan/SKILL.md +28 -1
- package/{commands/portfolio.md → skills/portfolio/SKILL.md} +8 -2
- package/skills/reconcile/SKILL.md +21 -0
- package/{commands/release.md → skills/release/SKILL.md} +16 -2
- package/skills/repo-audit/SKILL.md +7 -0
- package/skills/session-end/SKILL.md +13 -16
- package/skills/session-end/discovery-scan.md +1 -1
- package/skills/session-end/phase-3-6-tail.md +55 -9
- package/skills/session-end/plan-verification.md +2 -2
- package/skills/session-end/references/phase-5-issue-cleanup.md +9 -14
- package/skills/session-end/session-metrics-write.md +10 -0
- package/skills/session-plan/SKILL.md +18 -6
- package/skills/session-plan/references/session-plan-task-classification.md +2 -2
- package/skills/session-start/SKILL.md +5 -4
- package/skills/session-start/phase-8-5-express-path.md +6 -6
- package/skills/session-start/references/phase-1-5-session-continuity.md +1 -1
- package/skills/session-start/references/phase-2-7-portfolio-snapshot.md +1 -1
- package/skills/session-start/references/phase-4-ssot-environment-check.md +6 -4
- package/skills/spinout/SKILL.md +12 -1
- package/skills/sunset-review/SKILL.md +13 -0
- package/{commands/test.md → skills/test/SKILL.md} +10 -4
- package/skills/ux-grill/SKILL.md +20 -2
- package/skills/wave-executor/SKILL.md +14 -7
- package/skills/wave-executor/circuit-breaker.md +2 -0
- package/skills/wave-executor/references/wave-executor-state-init.md +18 -4
- package/skills/wave-executor/references/wave-loop-dispatch.md +5 -2
- package/skills/wave-executor/references/wave-loop-review.md +17 -1
- package/commands/autopilot.md +0 -80
- package/commands/bootstrap.md +0 -56
- package/commands/brainstorm.md +0 -48
- package/commands/debug.md +0 -36
- package/commands/discovery.md +0 -32
- package/commands/dispatcher.md +0 -59
- package/commands/eli5.md +0 -33
- package/commands/eval.md +0 -28
- package/commands/evolve.md +0 -10
- package/commands/grill.md +0 -45
- package/commands/memory-cleanup.md +0 -26
- package/commands/persona-panel.md +0 -121
- package/commands/plan.md +0 -15
- package/commands/reconcile.md +0 -23
- package/commands/repo-audit.md +0 -24
- package/commands/spinout.md +0 -15
- package/commands/sunset-review.md +0 -27
- package/commands/ux-grill.md +0 -51
|
@@ -10,6 +10,8 @@ description: >
|
|
|
10
10
|
portfolio. user: "/dispatcher" assistant: "Ranked 18 free repos — top recommendation: Pencil-Designs
|
|
11
11
|
(score 4.50, 90d stale). Confirm via the picker, I'll claim its lease atomically, then route you to
|
|
12
12
|
/session deep."</example>
|
|
13
|
+
user-invocable: true
|
|
14
|
+
argument-hint: "[--dry-run] [--repo <name>]"
|
|
13
15
|
model: sonnet
|
|
14
16
|
---
|
|
15
17
|
|
|
@@ -17,6 +19,25 @@ model: sonnet
|
|
|
17
19
|
|
|
18
20
|
> Cross-repo autopilot front-door — enumerate → rank → owner-AUQ → atomic claim → route. Read-only until the operator confirms; the only mutating step is the atomic `session.lock` claim, and it happens BEFORE any launch.
|
|
19
21
|
|
|
22
|
+
## Invocation
|
|
23
|
+
|
|
24
|
+
Invoked as `/dispatcher [--dry-run] [--repo <name>]` with arguments: **$ARGUMENTS**.
|
|
25
|
+
|
|
26
|
+
Parse `$ARGUMENTS` before doing anything else. The recognized flags are exactly those of the CLI table below and are passed straight through to `scripts/lib/dispatcher/cli.mjs`:
|
|
27
|
+
|
|
28
|
+
- `--dry-run` — run the non-mutating rank only; print the recommendation and the free-candidate table; do NOT claim any lease (skip Phase 3).
|
|
29
|
+
- `--repo <name>` — limit the human-readable output to a single `repoName` (informational; does not change ranking).
|
|
30
|
+
- `--start-dir <path>` — override the scan root (defaults to the confinement root).
|
|
31
|
+
- `--json` — emit the full `{ candidates, free, ranked, warnings, recommended }` object to stdout.
|
|
32
|
+
|
|
33
|
+
If `$ARGUMENTS` contains an unrecognized flag (starts with `--` but is not one of the above), inform the user:
|
|
34
|
+
|
|
35
|
+
```
|
|
36
|
+
Unknown flag '<flag>'. Recognized flags: --dry-run, --repo <name>, --start-dir <path>, --json.
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
Then continue with the remaining valid arguments.
|
|
40
|
+
|
|
20
41
|
## Soul
|
|
21
42
|
|
|
22
43
|
The dispatcher answers one question: *"of all my repos, which is the most worthwhile to work on right now, and is it free?"* It scans the confinement-root children, resolves each repo's free/busy status from its `session.lock` v2 lease (same lease semantics as the vault-status board), ranks only the FREE ones by `priority × staleness × readiness`, and recommends the single best one. You confirm via a picker, it claims the lease atomically (winning the race or excluding-and-re-ranking on a loss), then routes you to the entry command for that repo. Busy repos are listed-as-such, never selected.
|
|
@@ -119,13 +140,23 @@ Do NOT reinvent the claim — always go through `claimRepo`/`acquire`. The `ok:f
|
|
|
119
140
|
|
|
120
141
|
## Phase 4: Route
|
|
121
142
|
|
|
122
|
-
With the lease held,
|
|
123
|
-
|
|
124
|
-
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
143
|
+
With the lease held, routing splits on whether the chosen entry is model-invocable:
|
|
144
|
+
|
|
145
|
+
- **`/session housekeeping` or `/session deep`** (execution modes) and **`/plan`** (read-only
|
|
146
|
+
planning precursor) are NOT something the coordinator can invoke itself. `commands/session.md`
|
|
147
|
+
documents `session` and `plan` as **reserved terminal-only built-in names** — under
|
|
148
|
+
non-interactive (`claude -p`) invocation the bare form answers `"isn't available in this
|
|
149
|
+
environment"` — and `skills/plan/SKILL.md` additionally carries
|
|
150
|
+
`disable-model-invocation: true`, which blocks the `Skill` tool from invoking it regardless of
|
|
151
|
+
interactivity. The coordinator therefore **hands the operator the exact command to type**, one
|
|
152
|
+
line per option — `/session-orchestrator:session <mode>` or `/session-orchestrator:plan
|
|
153
|
+
[new|feature|retro]` — rather than attempting to invoke either itself.
|
|
154
|
+
- **`/discovery`** — read-only investigation precursor (maps scope; does not execute).
|
|
155
|
+
`skills/discovery/SKILL.md` carries no `disable-model-invocation` flag, so the coordinator MAY
|
|
156
|
+
invoke it directly via the `Skill` tool when the operator picks this option — no hand-off line
|
|
157
|
+
needed.
|
|
158
|
+
|
|
159
|
+
`/plan` and `/discovery` remain **read-only precursors**, NOT execution modes — the menu may route to them, but they only produce artifacts for a later execution session. The full mode taxonomy lives in the mode-selector surface (P2 of this epic); the dispatcher only routes to the entry command the operator picked.
|
|
129
160
|
|
|
130
161
|
## Phase 5: Edge cases
|
|
131
162
|
|
package/skills/eli5/SKILL.md
CHANGED
|
@@ -1,6 +1,8 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: eli5
|
|
3
3
|
description: Explain a topic like I'm a 5 year old — restate my last output, or a named topic, in plain words without dropping a single fact. Use when the user types /eli5 [topic], or says an answer was too technical, too long, or unclear about what he now has to do.
|
|
4
|
+
user-invocable: true
|
|
5
|
+
argument-hint: "[topic]"
|
|
4
6
|
model: inherit
|
|
5
7
|
tools: Read, Grep, Glob, Bash
|
|
6
8
|
---
|
|
@@ -9,6 +11,15 @@ tools: Read, Grep, Glob, Bash
|
|
|
9
11
|
|
|
10
12
|
Say it again in plain words. Same facts, in the order he needs them.
|
|
11
13
|
|
|
14
|
+
## Invocation
|
|
15
|
+
|
|
16
|
+
Invoked as `/eli5 [topic]` with arguments: **$ARGUMENTS**.
|
|
17
|
+
|
|
18
|
+
The argument is optional and is a topic, in prose. If `$ARGUMENTS` is empty, the target is my own last substantial output in this conversation; if there is none yet, say so rather than picking a topic for him.
|
|
19
|
+
|
|
20
|
+
- `/eli5` — restate what I just said.
|
|
21
|
+
- `/eli5 warum ist der Regel-Korpus voll?` — explain that, grounded in what this session measured.
|
|
22
|
+
|
|
12
23
|
## The frame
|
|
13
24
|
|
|
14
25
|
**Write for someone who knows this project but has not seen what you just saw.**
|
package/skills/eval/SKILL.md
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: eval
|
|
3
3
|
user-invocable: true
|
|
4
|
+
argument-hint: "[--session <id>] [--no-write] [--verify <run-id>]"
|
|
4
5
|
tags: [eval, measurement, quality, meta, standard]
|
|
5
6
|
model: sonnet
|
|
6
7
|
model-preference: sonnet
|
|
@@ -14,7 +15,7 @@ args-schema:
|
|
|
14
15
|
- flag: --verify
|
|
15
16
|
description: "Re-evaluate a stored run-id and diff per-dimension for scoring drift (exit 1 on drift)"
|
|
16
17
|
description: >
|
|
17
|
-
Use this skill to run an honest session-process evaluation (Standard v1, aiat-llm-eval/1.0) — score the last completed orchestrator session against the pre-registered rubric-
|
|
18
|
+
Use this skill to run an honest session-process evaluation (Standard v1, aiat-llm-eval/1.0) — score the last completed orchestrator session against the pre-registered rubric-v2 dimensions, run /eval, evaluate this session, produce an eval report, or re-verify a stored eval run for reproducibility. Deterministic-first with an optional advisory LLM judge; never produces a global score.
|
|
18
19
|
---
|
|
19
20
|
|
|
20
21
|
> **Platform Note:** State files use the platform's native directory: `.claude/` (Claude Code), `.codex/` (Codex CLI), or `.cursor/` (Cursor IDE). Shared metrics + the eval journal live in `.orchestrator/metrics/`. See `skills/_shared/platform-tools.md`.
|
|
@@ -22,14 +23,30 @@ description: >
|
|
|
22
23
|
# Eval Skill — Session-Process Evaluation (aiat-llm-eval/1.0)
|
|
23
24
|
|
|
24
25
|
On-demand, honest measurement of ONE completed orchestrator session against the
|
|
25
|
-
pre-registered **rubric-
|
|
26
|
+
pre-registered **rubric-v2** check set. The deterministic engine
|
|
26
27
|
(`scripts/eval-session.mjs` → `scripts/lib/eval/engine.mjs`) reads only local
|
|
27
|
-
metrics files (`sessions.jsonl` + `events.jsonl`), scores the
|
|
28
|
+
metrics files (`sessions.jsonl` + `events.jsonl`), scores the six deterministic
|
|
28
29
|
dimensions, appends a `session-eval` record to the journal, and optionally
|
|
29
|
-
renders an HTML report. An opt-in LLM judge overlays
|
|
30
|
+
renders an HTML report. An opt-in LLM judge overlays ONE advisory dimension
|
|
31
|
+
(`instruction-adherence`; `report-quality` was retired in rubric-v2, #1381).
|
|
30
32
|
|
|
31
33
|
The standard this skill implements is [`docs/eval/aiat-llm-eval-v1.md`](../../docs/eval/aiat-llm-eval-v1.md);
|
|
32
|
-
the frozen, content-hashed check set is [`skills/eval/rubric-
|
|
34
|
+
the frozen, content-hashed check set is [`skills/eval/rubric-v2.md`](./rubric-v2.md)
|
|
35
|
+
(stored records written before 2026-09-19 carry `rubric-v1` and are read against
|
|
36
|
+
[`rubric-v1.md`](./rubric-v1.md), which is never edited again).
|
|
37
|
+
|
|
38
|
+
## Invocation
|
|
39
|
+
|
|
40
|
+
Invoked as `/eval [--session <id>] [--no-write] [--verify <run-id>]` with arguments: **$ARGUMENTS** (parsed in Phase 1.2).
|
|
41
|
+
|
|
42
|
+
- `/eval` — evaluate the last completed session (resolution cascade), append the record, render the report.
|
|
43
|
+
- `/eval --session <id>` — evaluate a specific `session_id`.
|
|
44
|
+
- `/eval --no-write` — evaluate + report without appending to the journal (dry-run).
|
|
45
|
+
- `/eval --verify <run-id>` — re-score a stored run and diff for drift (the reproducibility proof; exit 1 on drift).
|
|
46
|
+
|
|
47
|
+
**On-demand `/eval` runs regardless of `eval.enabled`** — that flag gates only the automatic session-end eval phase (see Phase 1.1).
|
|
48
|
+
|
|
49
|
+
**Seams used:** `scripts/eval-session.mjs` (deterministic CLI) · `runEvalJudge` / `mergeJudgeDimensions` (`scripts/lib/eval/judge.mjs`, opt-in) · `writeEvalReport` (`scripts/lib/eval/report.mjs`) · `appendEvalRecord` (`scripts/lib/eval/sink.mjs`) · the `eval` config block · [`skills/eval/rubric-v2.md`](./rubric-v2.md) (frozen check set).
|
|
33
50
|
|
|
34
51
|
## Posture Contract (load-bearing — read before executing)
|
|
35
52
|
|
|
@@ -38,9 +55,9 @@ the frozen, content-hashed check set is [`skills/eval/rubric-v1.md`](./rubric-v1
|
|
|
38
55
|
- **Never guess.** Missing source data yields `cannot-determine` (a first-class,
|
|
39
56
|
non-error verdict) with an honest reason — never a fabricated `pass`/`fail`.
|
|
40
57
|
Do NOT "fill in" a missing KPI or infer a gate result the events do not show.
|
|
41
|
-
- **Deterministic before judge.** The
|
|
42
|
-
on their own. The judge (Phase 3) is opt-in, ADVISORY, and `uncalibrated`
|
|
43
|
-
|
|
58
|
+
- **Deterministic before judge.** The six deterministic dimensions are complete
|
|
59
|
+
on their own. The judge (Phase 3) is opt-in, ADVISORY, and `uncalibrated` —
|
|
60
|
+
never blend a judge verdict into the deterministic tally.
|
|
44
61
|
- **Journal is SSOT; the report is a derived view.** The append-only
|
|
45
62
|
`.orchestrator/metrics/eval.jsonl` is authoritative. The HTML report is
|
|
46
63
|
rebuildable from any stored record and is never authoritative over the journal.
|
|
@@ -134,10 +151,13 @@ node scripts/eval-session.mjs [--session <id>] --json \
|
|
|
134
151
|
On exit `1` (e.g. "no completed session found"), surface the message and stop —
|
|
135
152
|
do not retry with fabricated inputs.
|
|
136
153
|
|
|
137
|
-
Parse the emitted JSON record. It carries `dimensions[]` (
|
|
138
|
-
entries
|
|
139
|
-
|
|
140
|
-
|
|
154
|
+
Parse the emitted JSON record. It carries `dimensions[]` (6 deterministic
|
|
155
|
+
entries — the two reported-only ones, `guard-friction` and `efficiency-kpis`,
|
|
156
|
+
are always `not-applicable`), `kpis{}`, `provenance.rubric_sha256` (the sha256
|
|
157
|
+
of `rubric-v2.md`; `null` means the rubric file was not found and the append
|
|
158
|
+
will fail validation), `model`, `harness`, and `run_id`. Unless `--no-write` was
|
|
159
|
+
passed, the record is already appended to `.orchestrator/metrics/eval.jsonl` by
|
|
160
|
+
the CLI.
|
|
141
161
|
|
|
142
162
|
**Contamination check:** if the human-render/summary reports a peer-overlapped
|
|
143
163
|
window, note it — `verification-evidence` and `gate-health` will read
|
|
@@ -173,9 +193,13 @@ const merged = mergeJudgeDimensions(record, dimensions);
|
|
|
173
193
|
appendEvalRecord(merged, { path: '.orchestrator/metrics/eval.jsonl' });
|
|
174
194
|
```
|
|
175
195
|
|
|
176
|
-
-
|
|
177
|
-
"uncalibrated"` (the schema firewall rejects any other shape). Keep
|
|
178
|
-
visibly separated from the deterministic
|
|
196
|
+
- The judge dimension arrives `advisory: true` + `calibration_status:
|
|
197
|
+
"uncalibrated"` (the schema firewall rejects any other shape). Keep it
|
|
198
|
+
visibly separated from the deterministic six in the summary.
|
|
199
|
+
- `runEvalJudge` pre-computes a `facts` block (`computeRecordFacts`) from the
|
|
200
|
+
deterministic evidence strings and hands it to the judge OUTSIDE the
|
|
201
|
+
untrusted-data fence — anything countable is counted in code, never inferred
|
|
202
|
+
by the model. Do not reimplement that here.
|
|
179
203
|
- If `runEvalJudge` returns a non-ok `status` (e.g. dispatch failed), keep the
|
|
180
204
|
deterministic record as-is and note the judge was unavailable — the
|
|
181
205
|
deterministic evaluation is complete without it.
|
|
@@ -212,20 +236,20 @@ const res = writeEvalReport(record, { generatedAt: new Date().toISOString() });
|
|
|
212
236
|
Emit a compact, honest per-dimension summary. Status lines only — no global score.
|
|
213
237
|
|
|
214
238
|
```
|
|
215
|
-
## /eval — <session_id> (self-evaluation, aiat-llm-eval/1.0 · rubric-
|
|
239
|
+
## /eval — <session_id> (self-evaluation, aiat-llm-eval/1.0 · rubric-v2 · n=1, no CI)
|
|
216
240
|
|
|
217
241
|
Deterministic:
|
|
218
242
|
verification-evidence PASS <one-line evidence>
|
|
219
243
|
plan-fidelity PASS completion_rate=1.0 (score)
|
|
220
244
|
gate-health PASS <one-line evidence>
|
|
221
|
-
process-safety PASS <one-line evidence + guard
|
|
245
|
+
process-safety PASS <one-line evidence + both guard disclosures>
|
|
246
|
+
guard-friction N/A (reported: blocked=… warned=… loop.warning=… <attribution>)
|
|
222
247
|
efficiency-kpis N/A (reported: duration=…s waves=… agents=… tok_in=… tok_out=… carryover=…)
|
|
223
248
|
|
|
224
249
|
Judge (advisory, uncalibrated) [only when eval.judge != off]:
|
|
225
|
-
instruction-adherence <verdict> advisory
|
|
226
|
-
report-quality <verdict> advisory
|
|
250
|
+
instruction-adherence <verdict> advisory (rule <n> applied)
|
|
227
251
|
|
|
228
|
-
cannot-determine: <k> of
|
|
252
|
+
cannot-determine: <k> of 6 deterministic dimensions (<reasons>)
|
|
229
253
|
Report: .orchestrator/eval/reports/<run_id>.html
|
|
230
254
|
Journal: .orchestrator/metrics/eval.jsonl (appended: <yes|--no-write>)
|
|
231
255
|
Re-verify: node scripts/eval-session.mjs --verify <run_id>
|
|
@@ -252,6 +276,11 @@ node scripts/eval-session.mjs --verify <run-id> --json
|
|
|
252
276
|
changed since the record was written — investigate, do not overwrite.
|
|
253
277
|
- `--verify` reproduces the stored model + timestamp verbatim (no env override),
|
|
254
278
|
so a MATCH is a real reproducibility proof of the scoring, not of model output.
|
|
279
|
+
- **A cross-version DRIFT is not a defect.** A stored `rubric-v1` record
|
|
280
|
+
re-scored by today's rubric-v2 engine necessarily differs on `process-safety`
|
|
281
|
+
and reports `present-in-fresh-only: guard-friction`. Read the record's
|
|
282
|
+
`rubric_version` before treating a diff as a regression (#1400 replaces that
|
|
283
|
+
report with an explicit version verdict).
|
|
255
284
|
|
|
256
285
|
---
|
|
257
286
|
|
|
@@ -268,9 +297,9 @@ platform. Only the judge phase needs harness-specific tooling.
|
|
|
268
297
|
`skills/_shared/platform-tools.md` § Agent Dispatch Pattern.
|
|
269
298
|
- **`harness.platform`** on the record is resolved from `$SO_PLATFORM`
|
|
270
299
|
(falls back to `claude-code`) inside the engine — no skill action needed.
|
|
271
|
-
- The deterministic
|
|
272
|
-
available on all platforms; the judge overlay is a Claude-Code-only
|
|
273
|
-
|
|
300
|
+
- The deterministic six dimensions + the HTML report + `--verify` are fully
|
|
301
|
+
available on all platforms; the judge overlay is a Claude-Code-only
|
|
302
|
+
enrichment.
|
|
274
303
|
|
|
275
304
|
---
|
|
276
305
|
|
package/skills/eval/rubric-v1.md
CHANGED
|
@@ -5,6 +5,7 @@
|
|
|
5
5
|
- **Conforms to standard:** `aiat-llm-eval/1.0` — see [`docs/eval/aiat-llm-eval-v1.md`](../../docs/eval/aiat-llm-eval-v1.md)
|
|
6
6
|
- **Reference engine:** [`scripts/lib/eval/engine.mjs`](../../scripts/lib/eval/engine.mjs) (the executable scorers this document mirrors verbatim)
|
|
7
7
|
- **Hash binding:** the sha256 of THIS FILE is written to every record's `provenance.rubric_sha256`.
|
|
8
|
+
- **Superseded for NEW records by [`rubric-v2.md`](./rubric-v2.md) (2026-09-19, issue #1037).** This file stays frozen and authoritative for every record already carrying `rubric_version: "rubric-v1"` — its formulas are the ones those verdicts were produced under and must not be edited. v2 changed exactly two things: `process-safety` no longer fails on `destructive_guard.blocked`, and the guard counts moved to a new reported-only `guard-friction` dimension. The reasoning and the measurement live in rubric-v2 § Änderungen gegenüber v1.
|
|
8
9
|
|
|
9
10
|
> **Pre-Registration (leading principle, standard §1.1).** The checks below are
|
|
10
11
|
> **fixed BEFORE the first scored run executes against this rubric**. This document
|