open-multi-agent-kit 1.0.0 → 1.2.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +52 -1
- package/README.md +2 -2
- package/dist/approval-api.d.ts +5 -0
- package/dist/approval-api.d.ts.map +1 -0
- package/dist/approval-api.js +5 -0
- package/dist/approval-api.js.map +1 -0
- package/dist/cli/help.d.ts.map +1 -1
- package/dist/cli/help.js +1 -0
- package/dist/cli/help.js.map +1 -1
- package/dist/cli/session-picker.d.ts +2 -2
- package/dist/cli/session-picker.d.ts.map +1 -1
- package/dist/cli/session-picker.js.map +1 -1
- package/dist/commands/verified-run-cli.d.ts.map +1 -1
- package/dist/commands/verified-run-cli.js +69 -22
- package/dist/commands/verified-run-cli.js.map +1 -1
- package/dist/commands/verified-run-signal.d.ts +3 -0
- package/dist/commands/verified-run-signal.d.ts.map +1 -0
- package/dist/commands/verified-run-signal.js +15 -0
- package/dist/commands/verified-run-signal.js.map +1 -0
- package/dist/coordination/awareness.d.ts +32 -0
- package/dist/coordination/awareness.d.ts.map +1 -0
- package/dist/coordination/awareness.js +55 -0
- package/dist/coordination/awareness.js.map +1 -0
- package/dist/coordination/broker.d.ts +79 -0
- package/dist/coordination/broker.d.ts.map +1 -0
- package/dist/coordination/broker.js +202 -0
- package/dist/coordination/broker.js.map +1 -0
- package/dist/coordination/index.d.ts +16 -0
- package/dist/coordination/index.d.ts.map +1 -0
- package/dist/coordination/index.js +16 -0
- package/dist/coordination/index.js.map +1 -0
- package/dist/coordination/integration.d.ts +59 -0
- package/dist/coordination/integration.d.ts.map +1 -0
- package/dist/coordination/integration.js +126 -0
- package/dist/coordination/integration.js.map +1 -0
- package/dist/coordination/operation.d.ts +107 -0
- package/dist/coordination/operation.d.ts.map +1 -0
- package/dist/coordination/operation.js +231 -0
- package/dist/coordination/operation.js.map +1 -0
- package/dist/coordination/resource.d.ts +33 -0
- package/dist/coordination/resource.d.ts.map +1 -0
- package/dist/coordination/resource.js +95 -0
- package/dist/coordination/resource.js.map +1 -0
- package/dist/coordination/session.d.ts +74 -0
- package/dist/coordination/session.d.ts.map +1 -0
- package/dist/coordination/session.js +143 -0
- package/dist/coordination/session.js.map +1 -0
- package/dist/coordination/types.d.ts +68 -0
- package/dist/coordination/types.d.ts.map +1 -0
- package/dist/coordination/types.js +26 -0
- package/dist/coordination/types.js.map +1 -0
- package/dist/core/adaptorch-bridge.d.ts +1 -1
- package/dist/core/adaptorch-bridge.d.ts.map +1 -1
- package/dist/core/adaptorch-bridge.js +1 -1
- package/dist/core/adaptorch-bridge.js.map +1 -1
- package/dist/core/agent-session-runtime.d.ts.map +1 -1
- package/dist/core/agent-session-runtime.js +16 -16
- package/dist/core/agent-session-runtime.js.map +1 -1
- package/dist/core/agent-session.d.ts.map +1 -1
- package/dist/core/agent-session.js +35 -15
- package/dist/core/agent-session.js.map +1 -1
- package/dist/core/canonical-json.d.ts +3 -0
- package/dist/core/canonical-json.d.ts.map +1 -0
- package/dist/core/canonical-json.js +20 -0
- package/dist/core/canonical-json.js.map +1 -0
- package/dist/core/compaction/compaction.d.ts +7 -3
- package/dist/core/compaction/compaction.d.ts.map +1 -1
- package/dist/core/compaction/compaction.js +14 -29
- package/dist/core/compaction/compaction.js.map +1 -1
- package/dist/core/compaction/control-state.d.ts +41 -0
- package/dist/core/compaction/control-state.d.ts.map +1 -0
- package/dist/core/compaction/control-state.js +85 -0
- package/dist/core/compaction/control-state.js.map +1 -0
- package/dist/core/compaction/fallback.d.ts +71 -0
- package/dist/core/compaction/fallback.d.ts.map +1 -0
- package/dist/core/compaction/fallback.js +118 -0
- package/dist/core/compaction/fallback.js.map +1 -0
- package/dist/core/compaction/index.d.ts +2 -0
- package/dist/core/compaction/index.d.ts.map +1 -1
- package/dist/core/compaction/index.js +2 -0
- package/dist/core/compaction/index.js.map +1 -1
- package/dist/core/compaction/model-failover.d.ts +21 -0
- package/dist/core/compaction/model-failover.d.ts.map +1 -0
- package/dist/core/compaction/model-failover.js +46 -0
- package/dist/core/compaction/model-failover.js.map +1 -0
- package/dist/core/compaction/session-summary.d.ts +21 -0
- package/dist/core/compaction/session-summary.d.ts.map +1 -0
- package/dist/core/compaction/session-summary.js +27 -0
- package/dist/core/compaction/session-summary.js.map +1 -0
- package/dist/core/context-budget-headroom-candidates.d.ts +16 -1
- package/dist/core/context-budget-headroom-candidates.d.ts.map +1 -1
- package/dist/core/context-budget-headroom-candidates.js +31 -13
- package/dist/core/context-budget-headroom-candidates.js.map +1 -1
- package/dist/core/context-budget-headroom-types.d.ts +7 -1
- package/dist/core/context-budget-headroom-types.d.ts.map +1 -1
- package/dist/core/context-budget-headroom-types.js +0 -1
- package/dist/core/context-budget-headroom-types.js.map +1 -1
- package/dist/core/context-budget-headroom.d.ts +10 -0
- package/dist/core/context-budget-headroom.d.ts.map +1 -1
- package/dist/core/context-budget-headroom.js +11 -2
- package/dist/core/context-budget-headroom.js.map +1 -1
- package/dist/core/context-budget-v2-global-pass.d.ts +21 -0
- package/dist/core/context-budget-v2-global-pass.d.ts.map +1 -0
- package/dist/core/context-budget-v2-global-pass.js +203 -0
- package/dist/core/context-budget-v2-global-pass.js.map +1 -0
- package/dist/core/context-budget-v2-planned-items.d.ts +12 -0
- package/dist/core/context-budget-v2-planned-items.d.ts.map +1 -1
- package/dist/core/context-budget-v2-planned-items.js +21 -2
- package/dist/core/context-budget-v2-planned-items.js.map +1 -1
- package/dist/core/context-budget-v2-planner.d.ts.map +1 -1
- package/dist/core/context-budget-v2-planner.js +18 -19
- package/dist/core/context-budget-v2-planner.js.map +1 -1
- package/dist/core/context-budget-v2-scoring.d.ts +7 -1
- package/dist/core/context-budget-v2-scoring.d.ts.map +1 -1
- package/dist/core/context-budget-v2-scoring.js.map +1 -1
- package/dist/core/context-budget-v2-selection.d.ts +3 -0
- package/dist/core/context-budget-v2-selection.d.ts.map +1 -1
- package/dist/core/context-budget-v2-selection.js +6 -2
- package/dist/core/context-budget-v2-selection.js.map +1 -1
- package/dist/core/context-budget-v2-types.d.ts +6 -1
- package/dist/core/context-budget-v2-types.d.ts.map +1 -1
- package/dist/core/context-budget-v2-types.js +6 -1
- package/dist/core/context-budget-v2-types.js.map +1 -1
- package/dist/core/export-html/template.js +14 -0
- package/dist/core/extensions/index.d.ts +1 -1
- package/dist/core/extensions/index.d.ts.map +1 -1
- package/dist/core/extensions/index.js.map +1 -1
- package/dist/core/extensions/session-lifecycle-types.d.ts +14 -0
- package/dist/core/extensions/session-lifecycle-types.d.ts.map +1 -0
- package/dist/core/extensions/session-lifecycle-types.js +2 -0
- package/dist/core/extensions/session-lifecycle-types.js.map +1 -0
- package/dist/core/extensions/types.d.ts +4 -9
- package/dist/core/extensions/types.d.ts.map +1 -1
- package/dist/core/extensions/types.js.map +1 -1
- package/dist/core/http-dispatcher.d.ts.map +1 -1
- package/dist/core/http-dispatcher.js +10 -1
- package/dist/core/http-dispatcher.js.map +1 -1
- package/dist/core/keybindings.d.ts.map +1 -1
- package/dist/core/keybindings.js +3 -6
- package/dist/core/keybindings.js.map +1 -1
- package/dist/core/legacy-resource-policy.d.ts +7 -0
- package/dist/core/legacy-resource-policy.d.ts.map +1 -0
- package/dist/core/legacy-resource-policy.js +9 -0
- package/dist/core/legacy-resource-policy.js.map +1 -0
- package/dist/core/mcp/config.d.ts.map +1 -1
- package/dist/core/mcp/config.js +17 -3
- package/dist/core/mcp/config.js.map +1 -1
- package/dist/core/mcp/connection-queue.d.ts +12 -0
- package/dist/core/mcp/connection-queue.d.ts.map +1 -0
- package/dist/core/mcp/connection-queue.js +46 -0
- package/dist/core/mcp/connection-queue.js.map +1 -0
- package/dist/core/mcp/manager-runtime.d.ts +55 -0
- package/dist/core/mcp/manager-runtime.d.ts.map +1 -0
- package/dist/core/mcp/manager-runtime.js +66 -0
- package/dist/core/mcp/manager-runtime.js.map +1 -0
- package/dist/core/mcp/manager.d.ts +12 -68
- package/dist/core/mcp/manager.d.ts.map +1 -1
- package/dist/core/mcp/manager.js +33 -113
- package/dist/core/mcp/manager.js.map +1 -1
- package/dist/core/mcp/protocol.d.ts.map +1 -1
- package/dist/core/mcp/protocol.js +70 -40
- package/dist/core/mcp/protocol.js.map +1 -1
- package/dist/core/package-manager.d.ts +4 -0
- package/dist/core/package-manager.d.ts.map +1 -1
- package/dist/core/package-manager.js +59 -129
- package/dist/core/package-manager.js.map +1 -1
- package/dist/core/package-resource-patterns.d.ts +7 -0
- package/dist/core/package-resource-patterns.d.ts.map +1 -0
- package/dist/core/package-resource-patterns.js +99 -0
- package/dist/core/package-resource-patterns.js.map +1 -0
- package/dist/core/provider-default-models.d.ts +1 -0
- package/dist/core/provider-default-models.d.ts.map +1 -1
- package/dist/core/provider-default-models.js +1 -0
- package/dist/core/provider-default-models.js.map +1 -1
- package/dist/core/provider-display-names.d.ts.map +1 -1
- package/dist/core/provider-display-names.js +1 -0
- package/dist/core/provider-display-names.js.map +1 -1
- package/dist/core/provider-resilience.d.ts +9 -0
- package/dist/core/provider-resilience.d.ts.map +1 -1
- package/dist/core/provider-resilience.js +15 -0
- package/dist/core/provider-resilience.js.map +1 -1
- package/dist/core/reasoning-router-resolver.d.ts +7 -0
- package/dist/core/reasoning-router-resolver.d.ts.map +1 -1
- package/dist/core/reasoning-router-resolver.js +21 -3
- package/dist/core/reasoning-router-resolver.js.map +1 -1
- package/dist/core/reasoning-router-v4.d.ts +13 -7
- package/dist/core/reasoning-router-v4.d.ts.map +1 -1
- package/dist/core/reasoning-router-v4.js +8 -5
- package/dist/core/reasoning-router-v4.js.map +1 -1
- package/dist/core/resource-loader.d.ts.map +1 -1
- package/dist/core/resource-loader.js +4 -8
- package/dist/core/resource-loader.js.map +1 -1
- package/dist/core/run-execution-api.d.ts +8 -1
- package/dist/core/run-execution-api.d.ts.map +1 -1
- package/dist/core/run-execution-api.js +4 -0
- package/dist/core/run-execution-api.js.map +1 -1
- package/dist/core/run-journal.d.ts +1 -2
- package/dist/core/run-journal.d.ts.map +1 -1
- package/dist/core/run-journal.js +2 -19
- package/dist/core/run-journal.js.map +1 -1
- package/dist/core/session-compaction-service.d.ts +15 -0
- package/dist/core/session-compaction-service.d.ts.map +1 -1
- package/dist/core/session-compaction-service.js +17 -4
- package/dist/core/session-compaction-service.js.map +1 -1
- package/dist/core/session-listing.d.ts +51 -0
- package/dist/core/session-listing.d.ts.map +1 -0
- package/dist/core/session-listing.js +111 -0
- package/dist/core/session-listing.js.map +1 -0
- package/dist/core/session-manager.d.ts +12 -37
- package/dist/core/session-manager.d.ts.map +1 -1
- package/dist/core/session-manager.js +12 -118
- package/dist/core/session-manager.js.map +1 -1
- package/dist/core/session-message-text.d.ts +5 -0
- package/dist/core/session-message-text.d.ts.map +1 -0
- package/dist/core/session-message-text.js +9 -0
- package/dist/core/session-message-text.js.map +1 -0
- package/dist/core/skills-catalog-cache.d.ts +2 -0
- package/dist/core/skills-catalog-cache.d.ts.map +1 -1
- package/dist/core/skills-catalog-cache.js +71 -17
- package/dist/core/skills-catalog-cache.js.map +1 -1
- package/dist/core/todo-runtime-state.d.ts +12 -0
- package/dist/core/todo-runtime-state.d.ts.map +1 -1
- package/dist/core/todo-runtime-state.js +23 -0
- package/dist/core/todo-runtime-state.js.map +1 -1
- package/dist/core/tools/read.d.ts.map +1 -1
- package/dist/core/tools/read.js +1 -1
- package/dist/core/tools/read.js.map +1 -1
- package/dist/core/verified-run/authority-clock.d.ts +3 -0
- package/dist/core/verified-run/authority-clock.d.ts.map +1 -0
- package/dist/core/verified-run/authority-clock.js +25 -0
- package/dist/core/verified-run/authority-clock.js.map +1 -0
- package/dist/core/verified-run/authority-errors.d.ts +10 -0
- package/dist/core/verified-run/authority-errors.d.ts.map +1 -0
- package/dist/core/verified-run/authority-errors.js +17 -0
- package/dist/core/verified-run/authority-errors.js.map +1 -0
- package/dist/core/verified-run/authority-events.d.ts +126 -0
- package/dist/core/verified-run/authority-events.d.ts.map +1 -0
- package/dist/core/verified-run/authority-events.js +482 -0
- package/dist/core/verified-run/authority-events.js.map +1 -0
- package/dist/core/verified-run/authority-journal.d.ts +27 -0
- package/dist/core/verified-run/authority-journal.d.ts.map +1 -0
- package/dist/core/verified-run/authority-journal.js +83 -0
- package/dist/core/verified-run/authority-journal.js.map +1 -0
- package/dist/core/verified-run/authority-meaning.d.ts +13 -0
- package/dist/core/verified-run/authority-meaning.d.ts.map +1 -0
- package/dist/core/verified-run/authority-meaning.js +40 -0
- package/dist/core/verified-run/authority-meaning.js.map +1 -0
- package/dist/core/verified-run/authority-runtime.d.ts +82 -0
- package/dist/core/verified-run/authority-runtime.d.ts.map +1 -0
- package/dist/core/verified-run/authority-runtime.js +169 -0
- package/dist/core/verified-run/authority-runtime.js.map +1 -0
- package/dist/core/verified-run/authority-store.d.ts +204 -0
- package/dist/core/verified-run/authority-store.d.ts.map +1 -0
- package/dist/core/verified-run/authority-store.js +852 -0
- package/dist/core/verified-run/authority-store.js.map +1 -0
- package/dist/core/verified-run/broker.d.ts +8 -1
- package/dist/core/verified-run/broker.d.ts.map +1 -1
- package/dist/core/verified-run/broker.js +94 -90
- package/dist/core/verified-run/broker.js.map +1 -1
- package/dist/core/verified-run/candidate-policy.d.ts +10 -0
- package/dist/core/verified-run/candidate-policy.d.ts.map +1 -0
- package/dist/core/verified-run/candidate-policy.js +33 -0
- package/dist/core/verified-run/candidate-policy.js.map +1 -0
- package/dist/core/verified-run/candidate.d.ts.map +1 -1
- package/dist/core/verified-run/candidate.js +5 -10
- package/dist/core/verified-run/candidate.js.map +1 -1
- package/dist/core/verified-run/coordinator.d.ts +35 -17
- package/dist/core/verified-run/coordinator.d.ts.map +1 -1
- package/dist/core/verified-run/coordinator.js +111 -57
- package/dist/core/verified-run/coordinator.js.map +1 -1
- package/dist/core/verified-run/dag-phase.d.ts.map +1 -1
- package/dist/core/verified-run/dag-phase.js +1 -1
- package/dist/core/verified-run/dag-phase.js.map +1 -1
- package/dist/core/verified-run/dag-recovery.d.ts +5 -1
- package/dist/core/verified-run/dag-recovery.d.ts.map +1 -1
- package/dist/core/verified-run/dag-recovery.js +10 -4
- package/dist/core/verified-run/dag-recovery.js.map +1 -1
- package/dist/core/verified-run/event-parser.d.ts.map +1 -1
- package/dist/core/verified-run/event-parser.js +33 -0
- package/dist/core/verified-run/event-parser.js.map +1 -1
- package/dist/core/verified-run/evidence-read-projection.d.ts +4 -0
- package/dist/core/verified-run/evidence-read-projection.d.ts.map +1 -0
- package/dist/core/verified-run/evidence-read-projection.js +27 -0
- package/dist/core/verified-run/evidence-read-projection.js.map +1 -0
- package/dist/core/verified-run/evidence.d.ts +12 -1
- package/dist/core/verified-run/evidence.d.ts.map +1 -1
- package/dist/core/verified-run/evidence.js +51 -2
- package/dist/core/verified-run/evidence.js.map +1 -1
- package/dist/core/verified-run/git-candidate-preflight.d.ts +4 -0
- package/dist/core/verified-run/git-candidate-preflight.d.ts.map +1 -0
- package/dist/core/verified-run/git-candidate-preflight.js +34 -0
- package/dist/core/verified-run/git-candidate-preflight.js.map +1 -0
- package/dist/core/verified-run/git-effect-supervisor.d.ts +24 -0
- package/dist/core/verified-run/git-effect-supervisor.d.ts.map +1 -0
- package/dist/core/verified-run/git-effect-supervisor.js +174 -0
- package/dist/core/verified-run/git-effect-supervisor.js.map +1 -0
- package/dist/core/verified-run/git-execution.d.ts +27 -0
- package/dist/core/verified-run/git-execution.d.ts.map +1 -0
- package/dist/core/verified-run/git-execution.js +118 -0
- package/dist/core/verified-run/git-execution.js.map +1 -0
- package/dist/core/verified-run/git-plumbing.d.ts +37 -0
- package/dist/core/verified-run/git-plumbing.d.ts.map +1 -0
- package/dist/core/verified-run/git-plumbing.js +171 -0
- package/dist/core/verified-run/git-plumbing.js.map +1 -0
- package/dist/core/verified-run/git-publication-worker.d.ts +2 -0
- package/dist/core/verified-run/git-publication-worker.d.ts.map +1 -0
- package/dist/core/verified-run/git-publication-worker.js +54 -0
- package/dist/core/verified-run/git-publication-worker.js.map +1 -0
- package/dist/core/verified-run/git-sandbox-layout.d.ts +4 -0
- package/dist/core/verified-run/git-sandbox-layout.d.ts.map +1 -0
- package/dist/core/verified-run/git-sandbox-layout.js +29 -0
- package/dist/core/verified-run/git-sandbox-layout.js.map +1 -0
- package/dist/core/verified-run/git-worker-protocol.d.ts +19 -0
- package/dist/core/verified-run/git-worker-protocol.d.ts.map +1 -0
- package/dist/core/verified-run/git-worker-protocol.js +109 -0
- package/dist/core/verified-run/git-worker-protocol.js.map +1 -0
- package/dist/core/verified-run/git-worker-runtime.d.ts +19 -0
- package/dist/core/verified-run/git-worker-runtime.d.ts.map +1 -0
- package/dist/core/verified-run/git-worker-runtime.js +74 -0
- package/dist/core/verified-run/git-worker-runtime.js.map +1 -0
- package/dist/core/verified-run/journal.d.ts +2 -3
- package/dist/core/verified-run/journal.d.ts.map +1 -1
- package/dist/core/verified-run/journal.js.map +1 -1
- package/dist/core/verified-run/namespace-identity.d.ts.map +1 -1
- package/dist/core/verified-run/namespace-identity.js +1 -1
- package/dist/core/verified-run/namespace-identity.js.map +1 -1
- package/dist/core/verified-run/owned-execution.d.ts +10 -1
- package/dist/core/verified-run/owned-execution.d.ts.map +1 -1
- package/dist/core/verified-run/owned-execution.js +124 -31
- package/dist/core/verified-run/owned-execution.js.map +1 -1
- package/dist/core/verified-run/phase-context.d.ts +2 -0
- package/dist/core/verified-run/phase-context.d.ts.map +1 -1
- package/dist/core/verified-run/phase-context.js.map +1 -1
- package/dist/core/verified-run/projection.d.ts.map +1 -1
- package/dist/core/verified-run/projection.js +51 -1
- package/dist/core/verified-run/projection.js.map +1 -1
- package/dist/core/verified-run/publish-preflight.d.ts +12 -0
- package/dist/core/verified-run/publish-preflight.d.ts.map +1 -0
- package/dist/core/verified-run/publish-preflight.js +38 -0
- package/dist/core/verified-run/publish-preflight.js.map +1 -0
- package/dist/core/verified-run/publish-start.d.ts +5 -0
- package/dist/core/verified-run/publish-start.d.ts.map +1 -0
- package/dist/core/verified-run/publish-start.js +7 -0
- package/dist/core/verified-run/publish-start.js.map +1 -0
- package/dist/core/verified-run/recovery-projection.d.ts.map +1 -1
- package/dist/core/verified-run/recovery-projection.js +5 -0
- package/dist/core/verified-run/recovery-projection.js.map +1 -1
- package/dist/core/verified-run/recovery.d.ts +5 -1
- package/dist/core/verified-run/recovery.d.ts.map +1 -1
- package/dist/core/verified-run/recovery.js +10 -4
- package/dist/core/verified-run/recovery.js.map +1 -1
- package/dist/core/verified-run/run-plan.d.ts +17 -0
- package/dist/core/verified-run/run-plan.d.ts.map +1 -0
- package/dist/core/verified-run/run-plan.js +18 -0
- package/dist/core/verified-run/run-plan.js.map +1 -0
- package/dist/core/verified-run/run-publish-owned.d.ts +6 -0
- package/dist/core/verified-run/run-publish-owned.d.ts.map +1 -0
- package/dist/core/verified-run/run-publish-owned.js +199 -0
- package/dist/core/verified-run/run-publish-owned.js.map +1 -0
- package/dist/core/verified-run/run-publish.d.ts +19 -0
- package/dist/core/verified-run/run-publish.d.ts.map +1 -0
- package/dist/core/verified-run/run-publish.js +210 -0
- package/dist/core/verified-run/run-publish.js.map +1 -0
- package/dist/core/verified-run/run-status-owned.d.ts +4 -0
- package/dist/core/verified-run/run-status-owned.d.ts.map +1 -0
- package/dist/core/verified-run/run-status-owned.js +14 -0
- package/dist/core/verified-run/run-status-owned.js.map +1 -0
- package/dist/core/verified-run/run-status.d.ts +133 -0
- package/dist/core/verified-run/run-status.d.ts.map +1 -0
- package/dist/core/verified-run/run-status.js +193 -0
- package/dist/core/verified-run/run-status.js.map +1 -0
- package/dist/core/verified-run/run-types.d.ts +31 -0
- package/dist/core/verified-run/run-types.d.ts.map +1 -1
- package/dist/core/verified-run/run-types.js.map +1 -1
- package/dist/core/verified-run/storage.d.ts.map +1 -1
- package/dist/core/verified-run/storage.js +1 -1
- package/dist/core/verified-run/storage.js.map +1 -1
- package/dist/core/verified-run/supervisor-adapter.d.ts +77 -0
- package/dist/core/verified-run/supervisor-adapter.d.ts.map +1 -0
- package/dist/core/verified-run/supervisor-adapter.js +164 -0
- package/dist/core/verified-run/supervisor-adapter.js.map +1 -0
- package/dist/core/verified-run/verification-phase.d.ts.map +1 -1
- package/dist/core/verified-run/verification-phase.js +1 -1
- package/dist/core/verified-run/verification-phase.js.map +1 -1
- package/dist/core/verified-run/writer-phase.d.ts.map +1 -1
- package/dist/core/verified-run/writer-phase.js +1 -0
- package/dist/core/verified-run/writer-phase.js.map +1 -1
- package/dist/core/verified-run/writer-recovery.d.ts +2 -0
- package/dist/core/verified-run/writer-recovery.d.ts.map +1 -1
- package/dist/core/verified-run/writer-recovery.js +1 -0
- package/dist/core/verified-run/writer-recovery.js.map +1 -1
- package/dist/core/workload-shard-execution-types.d.ts +4 -1
- package/dist/core/workload-shard-execution-types.d.ts.map +1 -1
- package/dist/core/workload-shard-execution-types.js.map +1 -1
- package/dist/core/workload-shard-executor.d.ts +2 -2
- package/dist/core/workload-shard-executor.d.ts.map +1 -1
- package/dist/core/workload-shard-executor.js +16 -37
- package/dist/core/workload-shard-executor.js.map +1 -1
- package/dist/core/workload-shard-frontier.d.ts +5 -0
- package/dist/core/workload-shard-frontier.d.ts.map +1 -0
- package/dist/core/workload-shard-frontier.js +47 -0
- package/dist/core/workload-shard-frontier.js.map +1 -0
- package/dist/core/workload-shard-runner.d.ts.map +1 -1
- package/dist/core/workload-shard-runner.js +28 -17
- package/dist/core/workload-shard-runner.js.map +1 -1
- package/dist/index.d.ts +4 -5
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +3 -4
- package/dist/index.js.map +1 -1
- package/dist/main.d.ts.map +1 -1
- package/dist/main.js +2 -1
- package/dist/main.js.map +1 -1
- package/dist/metacognition/bridge-result.d.ts +148 -0
- package/dist/metacognition/bridge-result.d.ts.map +1 -0
- package/dist/metacognition/bridge-result.js +46 -0
- package/dist/metacognition/bridge-result.js.map +1 -0
- package/dist/metacognition/calibration-selective.d.ts +78 -0
- package/dist/metacognition/calibration-selective.d.ts.map +1 -0
- package/dist/metacognition/calibration-selective.js +167 -0
- package/dist/metacognition/calibration-selective.js.map +1 -0
- package/dist/metacognition/calibration.d.ts +62 -0
- package/dist/metacognition/calibration.d.ts.map +1 -0
- package/dist/metacognition/calibration.js +122 -0
- package/dist/metacognition/calibration.js.map +1 -0
- package/dist/metacognition/checkpoint.d.ts +60 -0
- package/dist/metacognition/checkpoint.d.ts.map +1 -0
- package/dist/metacognition/checkpoint.js +57 -0
- package/dist/metacognition/checkpoint.js.map +1 -0
- package/dist/metacognition/context7.d.ts +48 -0
- package/dist/metacognition/context7.d.ts.map +1 -0
- package/dist/metacognition/context7.js +154 -0
- package/dist/metacognition/context7.js.map +1 -0
- package/dist/metacognition/decision.d.ts +44 -0
- package/dist/metacognition/decision.d.ts.map +1 -0
- package/dist/metacognition/decision.js +111 -0
- package/dist/metacognition/decision.js.map +1 -0
- package/dist/metacognition/evaluation.d.ts +73 -0
- package/dist/metacognition/evaluation.d.ts.map +1 -0
- package/dist/metacognition/evaluation.js +97 -0
- package/dist/metacognition/evaluation.js.map +1 -0
- package/dist/metacognition/index.d.ts +32 -0
- package/dist/metacognition/index.d.ts.map +1 -0
- package/dist/metacognition/index.js +32 -0
- package/dist/metacognition/index.js.map +1 -0
- package/dist/metacognition/knowledge-action.d.ts +43 -0
- package/dist/metacognition/knowledge-action.d.ts.map +1 -0
- package/dist/metacognition/knowledge-action.js +60 -0
- package/dist/metacognition/knowledge-action.js.map +1 -0
- package/dist/metacognition/knowledge.d.ts +84 -0
- package/dist/metacognition/knowledge.d.ts.map +1 -0
- package/dist/metacognition/knowledge.js +165 -0
- package/dist/metacognition/knowledge.js.map +1 -0
- package/dist/metacognition/obligations.d.ts +82 -0
- package/dist/metacognition/obligations.d.ts.map +1 -0
- package/dist/metacognition/obligations.js +139 -0
- package/dist/metacognition/obligations.js.map +1 -0
- package/dist/metacognition/observation-validity.d.ts +83 -0
- package/dist/metacognition/observation-validity.d.ts.map +1 -0
- package/dist/metacognition/observation-validity.js +118 -0
- package/dist/metacognition/observation-validity.js.map +1 -0
- package/dist/metacognition/observe.d.ts +45 -0
- package/dist/metacognition/observe.d.ts.map +1 -0
- package/dist/metacognition/observe.js +59 -0
- package/dist/metacognition/observe.js.map +1 -0
- package/dist/metacognition/policy.d.ts +53 -0
- package/dist/metacognition/policy.d.ts.map +1 -0
- package/dist/metacognition/policy.js +155 -0
- package/dist/metacognition/policy.js.map +1 -0
- package/dist/metacognition/predictions.d.ts +78 -0
- package/dist/metacognition/predictions.d.ts.map +1 -0
- package/dist/metacognition/predictions.js +181 -0
- package/dist/metacognition/predictions.js.map +1 -0
- package/dist/metacognition/retrieval.d.ts +53 -0
- package/dist/metacognition/retrieval.d.ts.map +1 -0
- package/dist/metacognition/retrieval.js +170 -0
- package/dist/metacognition/retrieval.js.map +1 -0
- package/dist/metacognition/risk.d.ts +48 -0
- package/dist/metacognition/risk.d.ts.map +1 -0
- package/dist/metacognition/risk.js +179 -0
- package/dist/metacognition/risk.js.map +1 -0
- package/dist/metacognition/route-economics.d.ts +95 -0
- package/dist/metacognition/route-economics.d.ts.map +1 -0
- package/dist/metacognition/route-economics.js +110 -0
- package/dist/metacognition/route-economics.js.map +1 -0
- package/dist/metacognition/runtime-bridge.d.ts +81 -0
- package/dist/metacognition/runtime-bridge.d.ts.map +1 -0
- package/dist/metacognition/runtime-bridge.js +458 -0
- package/dist/metacognition/runtime-bridge.js.map +1 -0
- package/dist/metacognition/skills.d.ts +42 -0
- package/dist/metacognition/skills.d.ts.map +1 -0
- package/dist/metacognition/skills.js +220 -0
- package/dist/metacognition/skills.js.map +1 -0
- package/dist/metacognition/state.d.ts +110 -0
- package/dist/metacognition/state.d.ts.map +1 -0
- package/dist/metacognition/state.js +52 -0
- package/dist/metacognition/state.js.map +1 -0
- package/dist/metacognition/validation.d.ts +19 -0
- package/dist/metacognition/validation.d.ts.map +1 -0
- package/dist/metacognition/validation.js +58 -0
- package/dist/metacognition/validation.js.map +1 -0
- package/dist/metacognition/verification.d.ts +93 -0
- package/dist/metacognition/verification.d.ts.map +1 -0
- package/dist/metacognition/verification.js +132 -0
- package/dist/metacognition/verification.js.map +1 -0
- package/dist/metacognition/verifier.d.ts +49 -0
- package/dist/metacognition/verifier.d.ts.map +1 -0
- package/dist/metacognition/verifier.js +88 -0
- package/dist/metacognition/verifier.js.map +1 -0
- package/dist/modes/interactive/components/assistant-message.d.ts +2 -0
- package/dist/modes/interactive/components/assistant-message.d.ts.map +1 -1
- package/dist/modes/interactive/components/assistant-message.js +42 -26
- package/dist/modes/interactive/components/assistant-message.js.map +1 -1
- package/dist/modes/interactive/components/chat-container.d.ts +7 -0
- package/dist/modes/interactive/components/chat-container.d.ts.map +1 -0
- package/dist/modes/interactive/components/chat-container.js +14 -0
- package/dist/modes/interactive/components/chat-container.js.map +1 -0
- package/dist/modes/interactive/components/dynamic-border.d.ts.map +1 -1
- package/dist/modes/interactive/components/dynamic-border.js +4 -2
- package/dist/modes/interactive/components/dynamic-border.js.map +1 -1
- package/dist/modes/interactive/components/login-dialog.d.ts.map +1 -1
- package/dist/modes/interactive/components/login-dialog.js +0 -1
- package/dist/modes/interactive/components/login-dialog.js.map +1 -1
- package/dist/modes/interactive/components/session-selector-async-search.d.ts +33 -0
- package/dist/modes/interactive/components/session-selector-async-search.d.ts.map +1 -0
- package/dist/modes/interactive/components/session-selector-async-search.js +109 -0
- package/dist/modes/interactive/components/session-selector-async-search.js.map +1 -0
- package/dist/modes/interactive/components/session-selector-loaders.d.ts +6 -0
- package/dist/modes/interactive/components/session-selector-loaders.d.ts.map +1 -0
- package/dist/modes/interactive/components/session-selector-loaders.js +9 -0
- package/dist/modes/interactive/components/session-selector-loaders.js.map +1 -0
- package/dist/modes/interactive/components/session-selector-search.d.ts +1 -1
- package/dist/modes/interactive/components/session-selector-search.d.ts.map +1 -1
- package/dist/modes/interactive/components/session-selector-search.js.map +1 -1
- package/dist/modes/interactive/components/session-selector-tree.d.ts +16 -0
- package/dist/modes/interactive/components/session-selector-tree.d.ts.map +1 -0
- package/dist/modes/interactive/components/session-selector-tree.js +46 -0
- package/dist/modes/interactive/components/session-selector-tree.js.map +1 -0
- package/dist/modes/interactive/components/session-selector.d.ts +6 -1
- package/dist/modes/interactive/components/session-selector.d.ts.map +1 -1
- package/dist/modes/interactive/components/session-selector.js +36 -76
- package/dist/modes/interactive/components/session-selector.js.map +1 -1
- package/dist/modes/interactive/components/tool-execution-images.d.ts +32 -0
- package/dist/modes/interactive/components/tool-execution-images.d.ts.map +1 -0
- package/dist/modes/interactive/components/tool-execution-images.js +144 -0
- package/dist/modes/interactive/components/tool-execution-images.js.map +1 -0
- package/dist/modes/interactive/components/tool-execution.d.ts +3 -4
- package/dist/modes/interactive/components/tool-execution.d.ts.map +1 -1
- package/dist/modes/interactive/components/tool-execution.js +22 -71
- package/dist/modes/interactive/components/tool-execution.js.map +1 -1
- package/dist/modes/interactive/interactive-mode.d.ts.map +1 -1
- package/dist/modes/interactive/interactive-mode.js +7 -9
- package/dist/modes/interactive/interactive-mode.js.map +1 -1
- package/dist/modes/interactive/tui-diagnostics.d.ts +8 -4
- package/dist/modes/interactive/tui-diagnostics.d.ts.map +1 -1
- package/dist/modes/interactive/tui-diagnostics.js +35 -1
- package/dist/modes/interactive/tui-diagnostics.js.map +1 -1
- package/dist/modes/rpc/rpc-mode.d.ts.map +1 -1
- package/dist/modes/rpc/rpc-mode.js +14 -17
- package/dist/modes/rpc/rpc-mode.js.map +1 -1
- package/dist/modes/rpc/rpc-response.d.ts +5 -0
- package/dist/modes/rpc/rpc-response.d.ts.map +1 -0
- package/dist/modes/rpc/rpc-response.js +17 -0
- package/dist/modes/rpc/rpc-response.js.map +1 -0
- package/dist/observation/identity.d.ts +14 -0
- package/dist/observation/identity.d.ts.map +1 -0
- package/dist/observation/identity.js +29 -0
- package/dist/observation/identity.js.map +1 -0
- package/dist/observation/index.d.ts +13 -0
- package/dist/observation/index.d.ts.map +1 -0
- package/dist/observation/index.js +13 -0
- package/dist/observation/index.js.map +1 -0
- package/dist/observation/observe-mode.d.ts +46 -0
- package/dist/observation/observe-mode.d.ts.map +1 -0
- package/dist/observation/observe-mode.js +83 -0
- package/dist/observation/observe-mode.js.map +1 -0
- package/dist/observation/store.d.ts +53 -0
- package/dist/observation/store.d.ts.map +1 -0
- package/dist/observation/store.js +126 -0
- package/dist/observation/store.js.map +1 -0
- package/dist/observation/types.d.ts +72 -0
- package/dist/observation/types.d.ts.map +1 -0
- package/dist/observation/types.js +9 -0
- package/dist/observation/types.js.map +1 -0
- package/dist/observation/view.d.ts +39 -0
- package/dist/observation/view.d.ts.map +1 -0
- package/dist/observation/view.js +173 -0
- package/dist/observation/view.js.map +1 -0
- package/dist/utils/exif-orientation.d.ts.map +1 -1
- package/dist/utils/exif-orientation.js +2 -3
- package/dist/utils/exif-orientation.js.map +1 -1
- package/dist/utils/tools-manager.d.ts.map +1 -1
- package/dist/utils/tools-manager.js +3 -5
- package/dist/utils/tools-manager.js.map +1 -1
- package/docs/adaptorch-preview-spec.md +7 -7
- package/docs/atomic-commit-planning.md +69 -0
- package/docs/compaction.md +26 -1
- package/docs/extensions.md +14 -0
- package/docs/mcp.md +22 -0
- package/docs/metacognition.md +130 -0
- package/docs/model-catalog-refresh.md +218 -1
- package/docs/models.md +1 -1
- package/docs/providers.md +53 -0
- package/docs/runtime-algorithms.md +156 -7
- package/docs/session-format.md +8 -0
- package/docs/sessions.md +19 -0
- package/docs/skills.md +19 -0
- package/docs/terminal-setup.md +15 -0
- package/docs/tui.md +41 -0
- package/docs/verified-run-safety.md +75 -0
- package/examples/extensions/custom-provider-anthropic/index.ts +1 -1
- package/examples/extensions/custom-provider-anthropic/package-lock.json +2 -2
- package/examples/extensions/custom-provider-anthropic/package.json +1 -1
- package/examples/extensions/custom-provider-gitlab-duo/package.json +1 -1
- package/examples/extensions/doom-overlay/wad-finder.ts +7 -3
- package/examples/extensions/gondolin/package-lock.json +2 -2
- package/examples/extensions/gondolin/package.json +1 -1
- package/examples/extensions/sandbox/package-lock.json +2 -2
- package/examples/extensions/sandbox/package.json +1 -1
- package/examples/extensions/with-deps/package-lock.json +2 -2
- package/examples/extensions/with-deps/package.json +1 -1
- package/npm-shrinkwrap.json +18 -18
- package/package.json +6 -8
|
@@ -0,0 +1,130 @@
|
|
|
1
|
+
# Metacognitive Control Kernel
|
|
2
|
+
|
|
3
|
+
`src/metacognition/` implements the observation / control / evaluation loop from
|
|
4
|
+
`docs/OMK_metacognitive_control_algorithms_2026-09-19.md`, on top of the
|
|
5
|
+
skill-and-knowledge control kernel from `docs/OMK_skill_knowledge_control_2026-09-19.zip`.
|
|
6
|
+
|
|
7
|
+
This is a decision-rules library, not a performance claim. It manages what the
|
|
8
|
+
agent *expects*, which *obligations* remain, whether *checks* can actually
|
|
9
|
+
detect the defects they cover, and when a *strategy* should switch — all as
|
|
10
|
+
host-owned structured state, never as model self-report.
|
|
11
|
+
|
|
12
|
+
## Modules
|
|
13
|
+
|
|
14
|
+
| Module | Spec | Purpose |
|
|
15
|
+
| --- | --- | --- |
|
|
16
|
+
| `knowledge.ts` | §13 core | Claim/evidence gap inspection (`inspectKnowledge`) |
|
|
17
|
+
| `knowledge-action.ts` | §13 core | Bounded next action (`nextKnowledgeAction`) |
|
|
18
|
+
| `skills.ts` | §13 core | Capability-coverage skill planning (`planSkills`) |
|
|
19
|
+
| `retrieval.ts` | §13 core | Local BM25 (`searchCorpus`), owned-promise `AcquisitionPool` |
|
|
20
|
+
| `context7.ts` | §13 core | Approved-egress Context7 GET adapter (`Context7Client`) |
|
|
21
|
+
| `runtime-bridge.ts` | §13 attach | ClaimGraph/ObservationNode → kernel inputs (`toMetaState`) |
|
|
22
|
+
| `evaluation.ts` | §11, §15 | DR offline estimator, experience records, completion metrics |
|
|
23
|
+
| `observe.ts` | §13 observe | Observation-mode diagnostics (`attachMetaDiagnostics`) |
|
|
24
|
+
| `obligations.ts` | A / §4 | Change-atom rules → candidate vs required obligations |
|
|
25
|
+
| `predictions.ts` | B / §5 | Pre-registered prediction ledger, Brier/surprise scoring |
|
|
26
|
+
| `decision.ts` | C / §6 | Finite Bayes risk + one-step VOI (`experimentValue`) |
|
|
27
|
+
| `verifier.ts` | D / §7 | Obligation-scoped negative-control evaluation (`evaluateVerifier`) |
|
|
28
|
+
| `calibration.ts` | E+G / §8, §10 | Condition-bucket Beta records + drift demotion |
|
|
29
|
+
| `state.ts` | §3, §9.5 | `MetaState` tuple and the three finish states |
|
|
30
|
+
| `policy.ts` | F / §9, §14 | Feasibility gating + priority-table action selection |
|
|
31
|
+
| `checkpoint.ts` | §9.4 | One ordered checkpoint evaluation (`checkpoint`) |
|
|
32
|
+
|
|
33
|
+
## Invariants enforced
|
|
34
|
+
|
|
35
|
+
- Required approvals and required checks are hard constraints, not optimization
|
|
36
|
+
inputs.
|
|
37
|
+
- Predictions are registered before outcomes and never overwritten; edits append.
|
|
38
|
+
- Model-produced obligations stay candidates until host facts promote them.
|
|
39
|
+
- Mutant counts exclude compile-broken, environment-failed, and equivalent
|
|
40
|
+
mutants; an empty denominator reports `unknown`, never a score.
|
|
41
|
+
- `max VOI <= 0` never implies verified completion; receipts must bind to the
|
|
42
|
+
current candidate hash.
|
|
43
|
+
- Environment failures are recorded separately from code failures — neither
|
|
44
|
+
inflates nor silently discards the other.
|
|
45
|
+
|
|
46
|
+
## Primitive hardening
|
|
47
|
+
|
|
48
|
+
The September 21, 2026 review's WP00 fixes harden the local library without
|
|
49
|
+
changing the default CLI workflow or enabling a supervisor.
|
|
50
|
+
|
|
51
|
+
### Contracts and compatibility
|
|
52
|
+
|
|
53
|
+
- Operations cannot return to observation or proposal after authorization or
|
|
54
|
+
dispatch. Replanning requires a new operation.
|
|
55
|
+
- Observations and authorization inputs are copied and frozen. Dispatch must
|
|
56
|
+
match the saved intent, approval, policy version and lease generation, even
|
|
57
|
+
when a replacement permit is internally consistent.
|
|
58
|
+
- Permit flags and postconditions must be booleans. Binding strings are bounded
|
|
59
|
+
and nonempty, and generations use canonical decimal `Sequence` values.
|
|
60
|
+
- Only unresolved dispatched outcomes can settle. Repeating the current outcome
|
|
61
|
+
is idempotent, including history. Applied and confirmed-failure outcomes cannot
|
|
62
|
+
be overwritten. Cancellation is still not proof of termination.
|
|
63
|
+
- The broker validates and snapshots claims before changing state. Malformed or
|
|
64
|
+
sparse arrays cannot acquire or start effects, or expire another reservation.
|
|
65
|
+
- Exact claim-set binding includes generation; resource conflicts still ignore
|
|
66
|
+
generation. A changed generation requires readmission before `start`.
|
|
67
|
+
- `claimKey` uses a JSON tuple including generation instead of NUL delimiters.
|
|
68
|
+
Consumers must recompute keys rather than mix old and new formats. This
|
|
69
|
+
in-memory kernel provides no durable key migration.
|
|
70
|
+
- Instance IDs reject NUL and lengths over 4096; canonical keys also reject
|
|
71
|
+
backslashes. Deadline sums must be safe integers, and sequence increments
|
|
72
|
+
remain within the existing 40-digit wire limit.
|
|
73
|
+
|
|
74
|
+
### Numerical behavior and tests
|
|
75
|
+
|
|
76
|
+
Temperature scaling shifts log probabilities before dividing by temperature, so
|
|
77
|
+
`Number.MIN_VALUE` preserves a nonzero maximum and finite ties. Log loss rejects
|
|
78
|
+
clips whose upper boundary rounds to one and uses `Math.log1p`.
|
|
79
|
+
|
|
80
|
+
Clopper-Pearson bounds invert the beta survival probability without rounding
|
|
81
|
+
`1 - alpha` to one; zero failures use the closed form. Incomplete beta rejects
|
|
82
|
+
invalid arguments, non-convergence and invalid results. Bonferroni adjustment
|
|
83
|
+
rejects comparison-count overflow and alpha underflow. These floating-point
|
|
84
|
+
calculations are not interval-arithmetic proofs or next-action guarantees.
|
|
85
|
+
|
|
86
|
+
From `packages/coding-agent`:
|
|
87
|
+
|
|
88
|
+
```bash
|
|
89
|
+
LIVE_E2E=0 node ../../node_modules/vitest/dist/cli.js --run test/primitive-hardening.test.ts test/primitive-boundaries.test.ts test/primitive-numeric-oracle.test.ts
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
The original 41 regression cases reproduced 28 failures before the six source
|
|
93
|
+
fixes and passed afterwards. The seeded broker model runs 64 seeds with 512
|
|
94
|
+
transitions each, branch-count assertions, and an independent component-overlap
|
|
95
|
+
oracle. Additional sparse-array regressions cover a gap in the supplied patch.
|
|
96
|
+
These are regression cases, not independent product failures or task successes.
|
|
97
|
+
|
|
98
|
+
`test/fixtures/primitive-numeric-oracle.json` contains 282 fixed values from
|
|
99
|
+
SciPy 1.17.0 `scipy.stats.beta.isf(alpha, k + 1, n - k)` (one when `k == n`).
|
|
100
|
+
The grid uses `n` in `{1, 2, 10, 100, 300, 1000, 10000, 100000, 1000000}`, valid
|
|
101
|
+
unique `k` in `{0, 1, 2, n // 2, n - 1, n}`, and alpha in
|
|
102
|
+
`{0.1, 0.05, 0.01, 1e-6, 1e-12, 1e-20}`. Tolerance is
|
|
103
|
+
`1e-9 + 1e-8 * abs(expected)`. Before hardening, 48 values exceeded tolerance;
|
|
104
|
+
afterwards all passed, with maximum absolute error about `6.82e-14` on this grid.
|
|
105
|
+
Tests require neither Python nor network access. Other numerical domains are
|
|
106
|
+
not established by this fixture.
|
|
107
|
+
|
|
108
|
+
### Integration boundary
|
|
109
|
+
|
|
110
|
+
The coordination and metacognition barrels are re-exported from `src/index.ts`.
|
|
111
|
+
In the reviewed source, `AdmissionBroker`, `OperationLifecycle`, temperature
|
|
112
|
+
scaling and Clopper-Pearson bounds have no production execution call sites.
|
|
113
|
+
`AdmissionBroker.start` uses the strengthened `sameClaimSet`; tests exercise
|
|
114
|
+
real implementations, not a mock supervisor. Library availability does not
|
|
115
|
+
mean the default agent loop enforces these contracts.
|
|
116
|
+
|
|
117
|
+
The review's WP01-WP07 remain separate work: an end-to-end verified runtime path,
|
|
118
|
+
worker/descendant termination, durable recovery, immutable candidate publication,
|
|
119
|
+
loss-aware bridge mapping, shared UI/SDK states, and same-budget efficacy
|
|
120
|
+
experiments. Additional primitive proposals are not part of these six fixes.
|
|
121
|
+
No release, OS fencing, durable recovery or automatic completion-verification
|
|
122
|
+
claim follows from these tests. See [Development](development.md) for the
|
|
123
|
+
repository check and explicit test-filter workflow.
|
|
124
|
+
|
|
125
|
+
## Tests
|
|
126
|
+
|
|
127
|
+
`test/metacognition-*.test.ts` — 94 tests, including the 13 numerical checks
|
|
128
|
+
ported from `check_examples.py` (Brier 0.9025 at p=0.95 failure, VOI 0.84,
|
|
129
|
+
XOR two-bit synergy 0.40, probability-vector rejection) and the reference
|
|
130
|
+
kernel's contract tests.
|
|
@@ -1,8 +1,126 @@
|
|
|
1
1
|
# 모델 목록·thinking 갱신 기록
|
|
2
2
|
|
|
3
|
-
확인일: 2026-09-
|
|
3
|
+
확인일: 2026-09-22. 생성기와 공급자 어댑터를 수정한 뒤 `npm run models:refresh`로
|
|
4
4
|
두 카탈로그를 재생성했다. 생성 파일을 손으로 수정하지 않았다.
|
|
5
5
|
|
|
6
|
+
## 2026-09-22 갱신: Grok 4.7·MiMo v2.6과 reasoning/context 대조
|
|
7
|
+
|
|
8
|
+
`npm run models:refresh`를 종료 0으로 재생성했다. 전 소스가 응답했고 `--allow-partial`은 쓰지 않았다.
|
|
9
|
+
키 없는 Zyloo는 정적 6개를 유지했다. 이미지 카탈로그는 변화 없다.
|
|
10
|
+
|
|
11
|
+
| 항목 | 값 |
|
|
12
|
+
| --- | --- |
|
|
13
|
+
| 공급자 / 모델 | 40 / 1,826 → 1,847 |
|
|
14
|
+
| 추가 | 32 |
|
|
15
|
+
| 제거 | 11 |
|
|
16
|
+
| context 또는 maxTokens 변경 | 15 |
|
|
17
|
+
| thinking 맵 변경(기존 id) | 0 |
|
|
18
|
+
|
|
19
|
+
### 공식 원천과 대조
|
|
20
|
+
|
|
21
|
+
AdaptOrch `TopologyRouter`는 이 검증 DAG을 `hybrid`로 권고했다(width 2 exact, critical depth 4,
|
|
22
|
+
coupling density 0.66). 이는 실행 순서가 아니라 권고이다. 실제 검증은 공개 문서·공개 카탈로그 숫자와
|
|
23
|
+
생성 결과의 대조이며, 공급자 추론 호출은 하지 않았다.
|
|
24
|
+
|
|
25
|
+
| 모델 | 문서 context | 문서 effort | 카탈로그 |
|
|
26
|
+
| --- | --- | --- | --- |
|
|
27
|
+
| `grok-4.7` | 500,000 | `low/medium/high/xhigh`, 기본 `high`, 끄기 불가 | `xai` context 500,000, maxTokens 500,000. 네이티브는 `applyGrokThinking`이 `xhigh`까지 노출하고 `max/ultra`는 `xhigh` 별칭 |
|
|
28
|
+
| `grok-4.6` | 500,000 | 같은 사다리 | 변경 없음 |
|
|
29
|
+
| `grok-4.5` | 500,000 | `low/medium/high` (`xhigh` 없음) | `xhigh` 미노출 유지. 새 floor가 4.20 스냅샷에는 적용되지 않음 |
|
|
30
|
+
| MiMo v2.6 Pro/Flash | OpenRouter context 1,048,576, max completion 131,072 | route가 effort 목록을 선언하지 않음 | 선언 없는 사다리를 지어내지 않음. `reasoning: true`, map 없음 |
|
|
31
|
+
| GPT-5.6 계열 | OpenAI 문서 1.05M | `none/low/medium/high/xhigh/max` | 기존 생성기가 1,000,000으로 고정. 라이브 90개 모두 1,000,000 |
|
|
32
|
+
|
|
33
|
+
근거: [Grok 4.7](https://docs.x.ai/developers/models/grok-4.7),
|
|
34
|
+
[xAI reasoning](https://docs.x.ai/developers/model-capabilities/text/reasoning),
|
|
35
|
+
[OpenRouter models](https://openrouter.ai/api/v1/models),
|
|
36
|
+
[models.dev](https://models.dev/api.json).
|
|
37
|
+
|
|
38
|
+
### 생성기 보정
|
|
39
|
+
|
|
40
|
+
models.dev가 `grok-4.7`의 `reasoning_options`에 `xhigh`를 선언한다. 카탈로그 지연 시
|
|
41
|
+
`/thinking xhigh`가 `high`로 좁히지 않도록, xAI 문서의 "4.6 이후" 규칙을
|
|
42
|
+
`isDocumentedGrokXhighModel`로 두었다. `grok-4.20-*` 날짜 스냅샷은 이 사다리가 아니다.
|
|
43
|
+
|
|
44
|
+
### 추가·제거
|
|
45
|
+
|
|
46
|
+
추가는 `xai`·`openrouter`·`vercel-ai-gateway`·`github-copilot`·`opencode-go`의 `grok-4.7`,
|
|
47
|
+
Xiaomi 직접·토큰 플랜 3곳·OpenRouter·Vercel·OpenCode Go의 MiMo v2.6 Flash/Pro(와 Pro UltraSpeed),
|
|
48
|
+
OpenRouter `nex-agi/nex-n2.5-pro`·`mistralai/mistral-small-3.1-24b-instruct`,
|
|
49
|
+
Vercel `mixedbread/toast-1`·`quiverai/arrow-2`·`arrow-2-telos`, Hugging Face `tencent/Hy4-preview`다.
|
|
50
|
+
|
|
51
|
+
제거는 갱신 소스에서 더 이상 선정되지 않은 항목이다. NVIDIA `deepseek-v4-flash-0731`,
|
|
52
|
+
OpenCode `mimo-v2.5-free`, OpenRouter `anthropic/claude-opus-4`와 batch 별칭 7개,
|
|
53
|
+
`kwaipilot/kat-coder-pro-v2`가 해당한다. 공급자 폐기 공지나 모든 계정의 사용 불가는 아니다.
|
|
54
|
+
|
|
55
|
+
OpenRouter가 선언한 context·출력 상한 변경 15건은 목록 값을 그대로 반영했다.
|
|
56
|
+
예를 들면 `anthropic/claude-sonnet-4` context는 1,000,000에서 200,000으로, Aion 2.0/3.0 context는
|
|
57
|
+
131,072에서 1,048,576으로 바뀌었다. 이 숫자는 route 선언이지 공급자 원문 보증은 아니다.
|
|
58
|
+
|
|
59
|
+
### 검증과 한계
|
|
60
|
+
|
|
61
|
+
표적 vitest 12파일 151개 통과. `generate-models.ts` LSP diagnostics 없음.
|
|
62
|
+
공급자 추론, 전체 `npm run check`, build/install, commit/push는 실행하지 않았다.
|
|
63
|
+
계정별 사용 가능 여부와 실제 청구액은 검증 범위 밖이다.
|
|
64
|
+
|
|
65
|
+
## 2026-09-19 갱신: 라이브 재생성과 생성기 결함 3건 교정
|
|
66
|
+
|
|
67
|
+
`npm run models:refresh`를 종료0으로 재생성했다(전 소스 응답, `--allow-partial` 미사용,
|
|
68
|
+
키 없는 Zyloo는 정적 6개 유지). 결과는 **39 providers, 1,795→1,799 모델**, 추가 4·제거 0·
|
|
69
|
+
변경 22(가격 20, OpenRouter 선언 thinking 2), 이미지 카탈로그 54개 변화 없음.
|
|
70
|
+
|
|
71
|
+
첫 재생성에서는 제거 4·변경 220이 나왔고, 그중 상류 변화가 아닌 항목이 셋이었다.
|
|
72
|
+
각각 실패하는 검사를 먼저 쓰고(11개 RED) 생성기를 고친 뒤 다시 생성했다.
|
|
73
|
+
|
|
74
|
+
| 결함 | 원인 | 조치 |
|
|
75
|
+
| --- | --- | --- |
|
|
76
|
+
| `kimi-coding` 공급자 4개 전부 소실 | models.dev가 `kimi-for-coding` 키를 `kimi-code-plan-global`(api.kimi.ai)·`kimi-code-plan-cn`(api.kimi.com)으로 분리. 생성기는 옛 키만 읽어 조용히 빈 결과 | 세 키를 순서대로 조회. OMK endpoint(`api.kimi.com/coding`)·헤더·thinking 맵은 그대로. `kimi-coding-catalog.test.ts` |
|
|
77
|
+
| cursor 143·devin 55개 고정-노력 레인에 `xhigh/max` 추가 | 직전 커밋은 `--cursor-only`로 가족 단위 pass를 우회했지만 전체 재생성은 `applyModelMetadata`(Fable→xhigh/max, Opus 5→전체 사다리 등)를 정적 레인에도 적용 | devin/cursor 항목을 pass 이후에 붙여 fast path와 동일하게 유지. `fixed-effort-lanes.test.ts`가 정적 카탈로그와 생성 결과의 동일성을 검사 |
|
|
78
|
+
| OpenCode Zen `deepseek-v4.1-flash`가 구형 V4 맵(`high`, `xhigh→max`)만 받음 | V4.1 계약 보정 조건이 `opencode-go`만 인식 | [Zen 문서](https://opencode.ai/docs/zen/)가 같은 `chat/completions` gateway를 명시하므로 `opencode`도 `off/low/high/max`·`max_tokens`·`supportsReasoningEffort` 적용. `deepseek-v41-native.test.ts` route에 추가 |
|
|
79
|
+
|
|
80
|
+
두 번째 결함은 cursor/devin 항목이 정적 카탈로그와 다르게 저장될 때 즉시 실패하므로,
|
|
81
|
+
앞으로 fast path와 전체 재생성이 갈라지면 검사에서 드러난다.
|
|
82
|
+
|
|
83
|
+
### 추가된 항목
|
|
84
|
+
|
|
85
|
+
| 공급자 | 요청 ID | thinking | 가격(1M) | 근거 |
|
|
86
|
+
| --- | --- | --- | --- | --- |
|
|
87
|
+
| `opencode` | `deepseek-v4.1-flash` | `off/low/high/max`, `thinking.type`+`reasoning_effort`, `max_tokens` | $0.30/$1.20, cache $0.006 | Zen 문서 endpoint·가격표 |
|
|
88
|
+
| `opencode` | `qwen3.8-flash` | Messages 예산 경로(기존 Zen Qwen과 동일) | $0.15/$0.47, cache read $0.016·write $0.20 | Zen 문서 |
|
|
89
|
+
| `openrouter` | `z-ai/glm-5.3-flashx` | mandatory, `low/high/max` (route 선언) | $0.37/$1.25, cache $0.075 | OpenRouter `created` 09-18 |
|
|
90
|
+
| `openrouter` | `prism-ml/ternary-bonsai-2-27b` | optional off, `medium/xhigh` (route 선언) | $0.075/$0.50 | OpenRouter `created` 09-18 |
|
|
91
|
+
|
|
92
|
+
### 상류 변경(가격·선언)
|
|
93
|
+
|
|
94
|
+
- OpenCode Zen·Vercel의 `gpt-5.6-sol`은 "50% Off" 표시가 끝나 $4/$20(cache $0.40/$5)로 복귀. Vercel `gpt-5.6-sol-fast`는 $8/$40.
|
|
95
|
+
- OpenRouter DeepSeek 계열 인하: `deepseek-v4.1-flash` $0.15/$0.60, `deepseek-v4-pro` $0.54/$1.09, `-latest` 별칭 3개 동반 조정.
|
|
96
|
+
목록 값이며 DeepSeek 직접 API의 시간대별 가격과 다르다.
|
|
97
|
+
- OpenRouter `moonshotai/kimi-k3`·`~moonshotai/kimi-latest` $1.70/$8.50, `z-ai/glm-5.3` $0.91/$2.86, `meta/muse-glimmer-30b` $0.35/$1.50, Nemotron 3 Ultra·3.5 Lightning 출력 상한 상향.
|
|
98
|
+
- `upstage/solar-pro-3`(`off/minimal~high`)·`solar-pro4`(`off/minimal~max`)가 route 선언을 얻어 thinking 맵이 생겼다.
|
|
99
|
+
|
|
100
|
+
### 조사했으나 넣지 않은 항목
|
|
101
|
+
|
|
102
|
+
- **Amazon Bedrock Kimi K3** (09-18 GA). [모델 카드](https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-moonshot-ai-kimi-k3.html)
|
|
103
|
+
기준 ID `global.moonshotai.kimi-k3`(Global $3/$15, cache read $0.30·write $3.75) ·
|
|
104
|
+
`us.moonshotai.kimi-k3`(US $3.30/$16.50), in-region 없음, 1M context, 이미지 입력, Converse·도구·스트리밍 지원.
|
|
105
|
+
같은 카드가 "Converse는 이전 턴의 reasoning content가 포함되면 `InternalServerException`"을 명시하는데,
|
|
106
|
+
OMK `amazon-bedrock`은 non-Claude 모델의 thinking을 `reasoningContent`로 그대로 replay한다(`convertMessages`).
|
|
107
|
+
카탈로그만 넣으면 두 번째 턴부터 실패하는 경로를 광고하므로, 공급자에서 이전 턴 reasoning을 제거하는
|
|
108
|
+
수정과 유료 실검증을 묶은 후속 단위로 미룬다. models.dev bedrock 목록에도 아직 없다.
|
|
109
|
+
- **Qwen3.8-Omni-Flash** (Alibaba, 09-18): OpenRouter·Vercel·models.dev의 tool 지원 목록에 없다. Model Studio 직접 경로는 OMK 내장 공급자가 아니다.
|
|
110
|
+
- **Gemini 3.8 Live / Live Extended Thinking** (09-15~16): 오디오 네이티브 경로로 코딩 카탈로그 소스에 없다.
|
|
111
|
+
- **GPT-6 Astra Law** (`gpt-6-astra-law`): OpenAI가 "coming soon"으로만 공지, API 미제공.
|
|
112
|
+
- **Union Alpha**: `stealth/union-alpha`는 09-18 갱신에서 이미 빠졌고, 정식 `unbiased/pareto`($2.50/$7.50, cache $0.25, thinking 미선언)가 HEAD에 있다. 이번 갱신에서 변화 없음.
|
|
113
|
+
|
|
114
|
+
### 검증과 한계
|
|
115
|
+
|
|
116
|
+
표적 vitest 13파일 **149개 통과**(신규 `fixed-effort-lanes`·`kimi-coding-catalog`, 확장한
|
|
117
|
+
`deepseek-v41-native` 포함), 전체 `tsgo --noEmit`, 변경 파일 Biome, module-size·import-cycles
|
|
118
|
+
baseline, `git diff --check` 모두 종료0. 공급자 추론, 전체 `npm run check`, build/install,
|
|
119
|
+
commit/push는 실행하지 않았다. 계정별 사용 가능 여부와 실제 청구액은 검증 범위 밖이다.
|
|
120
|
+
|
|
121
|
+
변경 단위: `packages/ai/scripts/{generate-models,catalog-thinking}.ts`, 생성 카탈로그, 검사 3파일, 본 문서.
|
|
122
|
+
제안 메시지: `fix(ai): 모델 카탈로그 09-19 갱신과 생성기 결함 교정 (kimi-coding 소실·고정 레인 확장·Zen V4.1 계약)`.
|
|
123
|
+
|
|
6
124
|
## 2026-09-17 후속: OpenCode Go DeepSeek V4.1 ID 변경
|
|
7
125
|
|
|
8
126
|
[OpenCode Go 공식 endpoint 목록](https://opencode.ai/docs/go/)의 현재 ID는
|
|
@@ -238,3 +356,102 @@ npm run check
|
|
|
238
356
|
[Gemini thinking](https://ai.google.dev/gemini-api/docs/generate-content/thinking),
|
|
239
357
|
[DeepSeek thinking](https://api-docs.deepseek.com/guides/thinking_mode),
|
|
240
358
|
[Model Studio DeepSeek](https://www.alibabacloud.com/help/en/model-studio/deepseek-api).
|
|
359
|
+
|
|
360
|
+
## 2026-09-23 대조: MiMo V2.6 커버리지와 xhigh
|
|
361
|
+
|
|
362
|
+
카탈로그를 다시 생성하지 않았다. 2026-09-22 갱신이 넣은 V2.6 항목을 라이브 소스와
|
|
363
|
+
대조하고, 빠지면 실패하는 검사를 추가했다. 공급자 추론 호출은 하지 않았다.
|
|
364
|
+
|
|
365
|
+
### 커버리지
|
|
366
|
+
|
|
367
|
+
`mimo-v2.6-pro`와 `mimo-v2.6-flash`를 올리는 현재 OMK 공급자는 모두 이미 갖고 있다.
|
|
368
|
+
|
|
369
|
+
| 공급자 | 상류 | 카탈로그 |
|
|
370
|
+
| `xiaomi`, 토큰 플랜 cn/ams/sgp | models.dev `xiaomi`가 Flash/Pro/Pro UltraSpeed | 네 곳 모두 세 모델. 토큰 플랜은 `xiaomi` 목록을 미러 |
|
|
371
|
+
| `openrouter` | 라이브 `/api/v1/models`의 `xiaomi/mimo-v2.6-*` 3개 | 그대로 |
|
|
372
|
+
| `vercel-ai-gateway` | 라이브 `ai-gateway.vercel.sh/v1/models` 3개 | 그대로 |
|
|
373
|
+
| `opencode-go` | Flash/Pro. UltraSpeed 없음 | Flash/Pro |
|
|
374
|
+
| `opencode` | Zen은 `mimo-v2.6-flash-free`만, context 200,000 | 그대로. 유료 Pro/Flash는 Zen 목록에 없음 |
|
|
375
|
+
| `huggingface` | 라우터는 V2.5/V2.5-Pro만 | V2.6 없음이 맞음 |
|
|
376
|
+
|
|
377
|
+
OMK에 없는 게이트웨이(nano-gpt, Kilo, CrossModel, EmpirioLabs, LLM Gateway, DevPass)는
|
|
378
|
+
이번 범위가 아니다.
|
|
379
|
+
|
|
380
|
+
### xhigh는 없다
|
|
381
|
+
|
|
382
|
+
MiMo V2.6은 xhigh를 지원하지 않는다. 상류가 선언하지 않은 사다리를 만들지 않았다.
|
|
383
|
+
|
|
384
|
+
- Xiaomi Chat Completions는 `thinking.type`의 `enabled`/`disabled`만 받는다.
|
|
385
|
+
[openai-api](https://mimo.mi.com/docs/en-US/api/chat/openai-api), 2026-09-22 갱신.
|
|
386
|
+
- Responses 호환 `reasoning.effort`는 `none`/`low`/`medium`/`high`다. `none`만 끄고
|
|
387
|
+
나머지는 동작이 같다. 문서 원문: "The reasoning intensity is not differentiated at this stage."
|
|
388
|
+
[responses](https://mimo.mi.com/docs/en-US/api/chat/responses)
|
|
389
|
+
- models.dev의 `xiaomi`, 토큰 플랜 3곳, `openrouter`는 `reasoning_options: [{type: "toggle"}]`다.
|
|
390
|
+
OpenRouter 라이브도 `supported_reasoning_parameters: null`이다.
|
|
391
|
+
- 라이브 Vercel 게이트웨이는 `xiaomi/mimo-v2.6-*`에 `none`/`minimal`/`low`/`medium`/`high`만
|
|
392
|
+
선언한다. models.dev `vercel` 항목의 `xhigh`/`max`는 이 조회와 맞지 않아 따르지 않았다.
|
|
393
|
+
|
|
394
|
+
xhigh가 보이는 곳은 V2.5 라우트(Hugging Face, 일부 애그리게이터)이거나 MiMo Code 문서의
|
|
395
|
+
OpenAI 변형 목록이다. V2.6 계약이 아니다.
|
|
396
|
+
|
|
397
|
+
### 검증
|
|
398
|
+
|
|
399
|
+
`packages/ai/test/mimo-v26-catalog.test.ts` 46개 통과(종료 0). 21개 항목이 존재하고
|
|
400
|
+
`xhigh`/`max`/`ultra`를 노출하지 않는지, Xiaomi 4곳은 `thinkingFormat: "deepseek"` 토글인지
|
|
401
|
+
고정한다. 공급자 추론, `npm run check`, build/install, commit은 하지 않았다.
|
|
402
|
+
|
|
403
|
+
## 2026-09-23 갱신: Claude Opus 5.5와 GPT-6 Sol
|
|
404
|
+
|
|
405
|
+
AdaptOrch `TopologyRouter`는 증거, 생성, 검사를 `sequential`로 권고했다(width 1, critical depth 3).
|
|
406
|
+
실행 순서가 아니라 권고이다. `npm run generate-models`는 전 소스가 응답해 종료 0으로 끝났다.
|
|
407
|
+
공급자 추론은 하지 않았다.
|
|
408
|
+
|
|
409
|
+
| 모델 | 공식 원천 | 카탈로그 |
|
|
410
|
+
| Claude Opus 5.5 | `claude-opus-5-5`, context 1,000,000, output 128,000, effort `low/medium/high/xhigh/max`, thinking 끄기 불가 | `anthropic`, Bedrock 6개 지역 식별자, OpenRouter `anthropic/claude-opus-5.5`, Vercel `anthropic/claude-opus-5.5` |
|
|
411
|
+
| GPT-6 Sol | OpenAI 모델 색인에 없음. OpenRouter는 `none`부터 `max`, Vercel 라이브는 `high`까지 | OpenRouter `openai/gpt-6-sol`, `openai/gpt-6-sol-pro`와 batch. Vercel `openai/gpt-6-sol`. `openai` 공급자에는 넣지 않음 |
|
|
412
|
+
|
|
413
|
+
Vertex의 `claude-opus-5-5@default`는 OMK `google-vertex`가 쓰는 Gemini API가 아니므로 넣지 않았다.
|
|
414
|
+
기존 문서는 Opus 4.7/4.8과 Fable만 `xhigh`와 `max`를 갖는다고 적었지만, Opus 5.5 공식 문서가 같은 사다리를 선언한다.
|
|
415
|
+
|
|
416
|
+
검사: `packages/ai/test/opus55-gpt6sol-catalog.test.ts` 15개, `catalog-thinking.test.ts` 23개,
|
|
417
|
+
`mimo-v26-catalog.test.ts` 46개 통과. `catalog-thinking.ts`와 `bedrock-thinking.ts` 주 LSP는 오류 없다.
|
|
418
|
+
`npm run check`, build/install, commit은 하지 않았다.
|
|
419
|
+
|
|
420
|
+
근거: [Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview),
|
|
421
|
+
[models.dev](https://models.dev/api.json), [OpenRouter models](https://openrouter.ai/api/v1/models),
|
|
422
|
+
[Vercel AI Gateway models](https://ai-gateway.vercel.sh/v1/models).
|
|
423
|
+
|
|
424
|
+
### 2026-09-23 후속 검증: 라이브 대조, 상류 드리프트, stale 검사 정리
|
|
425
|
+
|
|
426
|
+
AdaptOrch `TopologyRouter`는 증거/검사/보고 3노드를 `parallel`로 권고했다
|
|
427
|
+
(추가 비용 없음, 권고일 뿐 실행 순서 아님). 재생성으로 `npm run generate-models`
|
|
428
|
+
종료 0, 1,876개 route.
|
|
429
|
+
|
|
430
|
+
- **라이브 대조 일치**: OpenRouter(`openai/gpt-6-sol` 4종, effort `none`~`max`), Vercel
|
|
431
|
+
(`openai/gpt-6-sol`은 `none/low/medium/high`만 선언 — `anthropic-messages` 예산 경로라
|
|
432
|
+
thinkingLevelMap 생략이 일치), Anthropic 문서(1M context, 128K output, adaptive
|
|
433
|
+
thinking 상시, effort 5단계) 모두 카탈로그 값과 일치했다.
|
|
434
|
+
- **상류 드리프트 반영**: 이전 단락의 "OpenAI 모델 색인에 없음"은 stale해졌다.
|
|
435
|
+
models.dev가 `openai`/`azure`/`opencode`의 Responses route에 `gpt-6-sol`과
|
|
436
|
+
`gpt-6-luna`를 새로 올렸고 네 소스 모두 effort `none`~`max`를 선언한다.
|
|
437
|
+
`catalog-thinking.ts`에 `GPT6_SOL_LUNA_ID` 규칙을 추가해 해당 route에
|
|
438
|
+
`off:"none"`/`minimal:null`/`low`~`max` 맵을 입혔다(Vercel은 anthropic-messages
|
|
439
|
+
예산 경로, openai-codex는 별도 천장이라 제외). OpenCode의 `claude-opus-5-5`와
|
|
440
|
+
OpenRouter의 `qwen/qwen3.8-omni-flash`(effort 미선언, 일반 Qwen 규칙 적용)도
|
|
441
|
+
상류 신규 등재로 들어왔다.
|
|
442
|
+
- **stale 검사 정리**: `thinking-max-level.test.ts`가 기대하던 `z-ai/glm-5.2:batch`는
|
|
443
|
+
OpenRouter 라이브 목록에서 제거됐고(현재 `z-ai/glm-5.2`, `z-ai/glm-5.2:free`만 존재),
|
|
444
|
+
재생성 카탈로그는 정확하므로 검사 대상에서 빼고 주석을 남겼다.
|
|
445
|
+
- **정규식 대칭 보정**: `OPUS_55_ID`의 점 철자(`claude-opus-5.5`) 대안부가 `^|/`만
|
|
446
|
+
허용하던 것을 `^|[./]`로 맞췄다. 오늘 카탈로그에는 해당 형태가 없어 동작 변화는 없다.
|
|
447
|
+
- **회귀 가드 추가**: `opus55-gpt6sol-catalog.test.ts`에 `claude-opus-5` 본체가 여전히
|
|
448
|
+
`off`를 노출하는지, `openai`/`azure`/`opencode`의 `gpt-6-sol`이 선언 사다리를 노출하는지,
|
|
449
|
+
`gpt-6-sol-fast`를 지어내지 않는지 검사를 추가했다(20개로 확장).
|
|
450
|
+
- **기존 한계 기록**: `bedrock-thinking.ts`의 `opus-(?:4-[678]|5(?:-5)?)(?:-|$)`는
|
|
451
|
+
`opus-5-6` 같은 미래 ID도 매치한다(기존 `5` 분기와 동일한 관용, 이번 변경의 회귀 아님).
|
|
452
|
+
`grok-thinking.ts` 정적 테이블은 정확-일치 키만 지원해 `grok-4.8+` 변형이 들어오면
|
|
453
|
+
`isDocumentedGrokXhighModel` floor가 xhigh 맵은 입혀도 런타임 compat은 미적용이다.
|
|
454
|
+
|
|
455
|
+
재검증: 표적 vitest 14파일 204개 중 203 통과·1 skip(유료 Bedrock E2E). 변경 파일 Biome,
|
|
456
|
+
`tsgo --noEmit`, `git diff --check` 종료 0. 공급자 추론, `npm run check`, build/install,
|
|
457
|
+
commit은 하지 않았다. 계정별 사용 가능 여부와 실제 청구액은 검증 범위 밖이다.
|
package/docs/models.md
CHANGED
|
@@ -22,7 +22,7 @@ The NVIDIA catalog is filtered against NIM's live model IDs and known compatibil
|
|
|
22
22
|
limits. A historical GLM route is not evidence that NIM still lists it. Use the
|
|
23
23
|
current model selector rather than copying a removed ID from older examples.
|
|
24
24
|
|
|
25
|
-
The [2026-09-
|
|
25
|
+
The [2026-09-22 catalog refresh](model-catalog-refresh.md) records new models,
|
|
26
26
|
provider-specific thinking ladders, and verification limits. Model discovery does
|
|
27
27
|
not prove that an account can invoke that model.
|
|
28
28
|
|
package/docs/providers.md
CHANGED
|
@@ -271,6 +271,7 @@ omk
|
|
|
271
271
|
| Fireworks | `FIREWORKS_API_KEY` | `fireworks` |
|
|
272
272
|
| Together AI | `TOGETHER_API_KEY` | `together` |
|
|
273
273
|
| Kimi For Coding | `KIMI_API_KEY` | `kimi-coding` |
|
|
274
|
+
| WorkBuddy (CodeBuddy) | `WORKBUDDY_API_KEY` | `workbuddy` |
|
|
274
275
|
| Meta Model API | `META_API_KEY` (or `META_MODEL_API_KEY`, `MODEL_API_KEY`) | `meta` |
|
|
275
276
|
| MiniMax | `MINIMAX_API_KEY` | `minimax` |
|
|
276
277
|
| MiniMax (China) | `MINIMAX_CN_API_KEY` | `minimax-cn` |
|
|
@@ -282,6 +283,58 @@ omk
|
|
|
282
283
|
|
|
283
284
|
Reference for environment variables and `auth.json` keys: [`const envMap`](https://github.com/dmae97/omk/blob/main/packages/ai/src/env-api-keys.ts) in [`packages/ai/src/env-api-keys.ts`](https://github.com/dmae97/omk/blob/main/packages/ai/src/env-api-keys.ts).
|
|
284
285
|
|
|
286
|
+
#### WorkBuddy (Tencent CodeBuddy)
|
|
287
|
+
|
|
288
|
+
The cloud inference behind Tencent's CodeBuddy / WorkBuddy agents. Plan keys start
|
|
289
|
+
with `ck_` and are issued by the CodeBuddy login flow:
|
|
290
|
+
|
|
291
|
+
```bash
|
|
292
|
+
export WORKBUDDY_API_KEY=ck_...
|
|
293
|
+
omk --provider workbuddy --model glm-5.3
|
|
294
|
+
```
|
|
295
|
+
|
|
296
|
+
Endpoint and limits, verified against the live service on 2026-09-19:
|
|
297
|
+
|
|
298
|
+
- Chat requests go to `https://www.workbuddy.ai/v2/chat/completions` and must
|
|
299
|
+
stream; a non-streaming request is rejected. The same key is accepted by
|
|
300
|
+
`www.codebuddy.ai`, not by the `*.cn` hosts.
|
|
301
|
+
- The first message must be a `system` message. OMK sends the system prompt in
|
|
302
|
+
that position and prepends an empty one when a session has no system prompt.
|
|
303
|
+
- `max_tokens` carries the output cap; `max_completion_tokens` is ignored
|
|
304
|
+
upstream.
|
|
305
|
+
- Images are sent as data URIs, which is the only form the endpoint accepts.
|
|
306
|
+
- Thinking levels expose only the effort values the provider declares for that
|
|
307
|
+
model (`glm-5.3`: low/high/max, `glm-5.2`: high/xhigh, `gpt-5.6-*`:
|
|
308
|
+
low/medium/high/xhigh, …). The endpoint accepts `reasoning_effort`, but its
|
|
309
|
+
effect on reasoning volume is not verified, and no `off` value is offered
|
|
310
|
+
because none was verified. Thinking can therefore not be disabled per request.
|
|
311
|
+
- Model entitlement is per plan: the built-in catalog lists the 27 lanes verified
|
|
312
|
+
for one account, including eight the CLI's bundled product JSON never declares
|
|
313
|
+
(`claude-opus-5`, `claude-opus-4.6`, `claude-sonnet-4.6`,
|
|
314
|
+
`deepseek-v4.1-flash`, `gemini-3.8-flash`, `gpt-6-astra`, `grok-4.6`,
|
|
315
|
+
`hy4-preview`) which a cross-sweep of this repository's own model ids found.
|
|
316
|
+
`glm-5.0` answered `429` while its siblings answered `200`, so it is not
|
|
317
|
+
listed; the CLI's role aliases (`fast-model`, `balanced-model`,
|
|
318
|
+
`primary-model`, `deep-model`) resolve server-side to targets this catalog
|
|
319
|
+
cannot declare and are not listed either.
|
|
320
|
+
- Billing is metered in the provider's own credit unit, which OMK reports as
|
|
321
|
+
zero cost rather than inventing a USD rate. Measured per-call charges on one
|
|
322
|
+
account (small prompts): `claude-opus-5` 0.26, `claude-opus-4.6` 0.20,
|
|
323
|
+
`claude-sonnet-4.6` 0.12, `grok-4.6` 0.06–0.08, `gpt-5.6-sol` 0.04,
|
|
324
|
+
`gemini-3.8-flash` 0.02–0.06, `glm-5.3` 0.01, and 0 credits on `hy3`,
|
|
325
|
+
`deepseek-v3-0324`, `deepseek-v4.1-flash`, `glm-5.2`, `kimi-k2.5/k2.6` and
|
|
326
|
+
`minimax-m3`. A charge scales with tokens, so these are magnitudes, not
|
|
327
|
+
tariffs.
|
|
328
|
+
- Plan allowances (provider documentation, 2026-09-19): Free 100 credits/month,
|
|
329
|
+
Pro $10/month or $96/year for 2,000 credits/month, plus a limited-time
|
|
330
|
+
activity bonus of 30 credits/day on Free and 50/day on Pro. Credits are issued
|
|
331
|
+
monthly, do not carry over, and new accounts get 250 credits valid for 14 days;
|
|
332
|
+
Pro has a 7-day trial of 500 credits. Check the provider console for the
|
|
333
|
+
account's own balance: at the measured rates, the Free plan's 100 credits is
|
|
334
|
+
roughly 380 calls at `claude-opus-5` size, ~830 at `claude-sonnet-4.6`, and
|
|
335
|
+
1,600-5,000 at `gemini-3.8-flash`, while the 0-credit lanes above did not draw
|
|
336
|
+
on the allowance at all in these probes.
|
|
337
|
+
|
|
285
338
|
#### NVIDIA NIM
|
|
286
339
|
|
|
287
340
|
Set `NVIDIA_API_KEY` and select a currently listed NVIDIA model with `/model`.
|
|
@@ -21,6 +21,9 @@ baseline below is history, not current truth.
|
|
|
21
21
|
| Provider retry/failover classification | yes | `provider-retry`, `session-failure-cause` | default | resilience/classification tests | released | not measured |
|
|
22
22
|
| MCP descriptor injection screen | yes | `mcp/manager` import path | default | quarantine tests | released | pattern rule score, not calibrated risk |
|
|
23
23
|
| verified-run coordinator + evidence | yes | verified-run paths | opt-in command | coordinator/evidence tests | working tree | scope-limited binding, not general correctness |
|
|
24
|
+
| Atomic commit planner (`planAtomicCommits`) | yes | public agent API only; no live commit caller | explicit function call | local planner/API/property tests, not a CI receipt | working tree | not measured |
|
|
25
|
+
|
|
26
|
+
Named tests are evidence locations, not a claim that this exact working tree passed remote CI. Source call paths and fresh local results must be checked before promoting a gate.
|
|
24
27
|
|
|
25
28
|
Gates are reported per row so a green "implemented" never upgrades itself to
|
|
26
29
|
"released" or "measured".
|
|
@@ -44,6 +47,77 @@ a release, or evidence of latency/quality improvement. Resource normalization,
|
|
|
44
47
|
slot-aware ranking, fairness, and equal-budget runtime comparisons remain separate
|
|
45
48
|
work.
|
|
46
49
|
|
|
50
|
+
## Final claim resolution is abort-bound (2026-09-19 audit F02)
|
|
51
|
+
|
|
52
|
+
The dag-v2 frontier re-resolves each prepared call's resource claims from its
|
|
53
|
+
exact post-hook arguments before admission. That wait was a bare `await`; the
|
|
54
|
+
initial scheduling pass in `tool-dag-memo` was already raced against the run's
|
|
55
|
+
abort signal, so an extension `resourceClaims()` that never settled on the
|
|
56
|
+
second call pinned a cancelled run until the callback returned on its own.
|
|
57
|
+
`resolveFinalResolution` now uses the same `awaitWithAbort` boundary. The
|
|
58
|
+
callback itself is not killed — a plain Promise cannot be — but the aborted
|
|
59
|
+
outcome settles the call, the frontier loop exits on the same signal, and a late
|
|
60
|
+
fulfilment or rejection lands in a finished batch and admits nothing. In-flight
|
|
61
|
+
peers keep the existing abort contract (`aborted`, `executionStarted: true`),
|
|
62
|
+
and the authorization hook still runs once. Coverage:
|
|
63
|
+
`packages/agent/test/tool-dag-final-claims-abort.test.ts`. This is the wait
|
|
64
|
+
boundary only; isolating a non-cooperative extension needs a killable execution
|
|
65
|
+
boundary, which this change does not add.
|
|
66
|
+
|
|
67
|
+
## Deferred claim freshness and cancelled memo requests (2026-09-22)
|
|
68
|
+
|
|
69
|
+
The live path is `agentLoop` → `schedulePlannedDagLevels` → `runDagFrontier`.
|
|
70
|
+
A deferred call retains its prepared arguments and one-time authorization, and
|
|
71
|
+
admission always re-resolves its claims: a cached resolution that survived the
|
|
72
|
+
wait could execute a stale scope (A holds X, B defers on X, C starts on Z, and
|
|
73
|
+
after A settles B runs on Z using the cached X claim). Re-resolution preserves
|
|
74
|
+
conflict exclusion against both earlier pending calls and later running calls,
|
|
75
|
+
and stays abort-bound, so a late claim result cannot admit an aborted call.
|
|
76
|
+
|
|
77
|
+
The cached resolution is still reused for one thing — proving the deferral.
|
|
78
|
+
A deferral decision is only "wait", so a stale conflict answer executes nothing,
|
|
79
|
+
while re-invoking an extension `resourceClaims()` callback on every scan of the
|
|
80
|
+
wait paid real I/O for the same answer. The resolver-budget regression shows a
|
|
81
|
+
deferred call blocked behind a slow peer invoking the callback 5 times when the
|
|
82
|
+
answer was re-resolved per scan, and 3 times with the split: the scheduling
|
|
83
|
+
resolution, one conflict-proving resolution, and the fresh resolution required
|
|
84
|
+
before admission. Claim callbacks may still run again after a wait, so they must
|
|
85
|
+
not be used as authorization side effects.
|
|
86
|
+
|
|
87
|
+
The per-run schedule memo checks cancellation before inspecting arguments or
|
|
88
|
+
serializing its key, including warm-cache requests. An already-aborted request
|
|
89
|
+
returns no plan and does not promote cache recency.
|
|
90
|
+
|
|
91
|
+
`assignDagDependencies` also returns an empty-edge graph directly when every
|
|
92
|
+
resolution is read-only (including empty claims). In the 128-entry regression,
|
|
93
|
+
resolved-entry visits fell from 16,256 to 128 with the same dependency graph.
|
|
94
|
+
Mixed writes and exclusive barriers still use the existing conflict scan. This
|
|
95
|
+
is an operation-count result for that fixture, not measured end-to-end latency,
|
|
96
|
+
model quality, CI certification or a release claim.
|
|
97
|
+
|
|
98
|
+
Coverage: `tool-dag-deferred-refresh.test.ts`,
|
|
99
|
+
`tool-dag-deferred-resolver-budget.test.ts`,
|
|
100
|
+
`tool-dag-memo-cancellation.test.ts`, and
|
|
101
|
+
`tool-dag-readonly-dependencies.test.ts` in `packages/agent/test/`.
|
|
102
|
+
The separate `planAtomicCommits` export remains a pure public API,
|
|
103
|
+
not a live automatic commit coordinator. ECRAF and automatic shards
|
|
104
|
+
are not activated by these changes.
|
|
105
|
+
|
|
106
|
+
## Reasoning router resolver contract (2026-09-19 audit F05/F06)
|
|
107
|
+
|
|
108
|
+
The low-confidence escalation in `resolveThinkingLevelV4WithUncertainty` is
|
|
109
|
+
monotone at a fixed bias and hint only; it is not a floor at the class base
|
|
110
|
+
level, because a negative learning bias is applied before the +1 step (`debug`
|
|
111
|
+
with `bias=-2` still resolves to `medium`). The docstring previously claimed the
|
|
112
|
+
stronger floor. A policy that must never drop below the base level needs an
|
|
113
|
+
explicit safety floor, which is a cost decision not taken here. The shared
|
|
114
|
+
resolver also normalizes `bias` and `escalationSteps` to finite integers
|
|
115
|
+
(non-finite → 0, fractions truncated toward zero) so a corrupted value keeps the
|
|
116
|
+
class's own level instead of walking the ladder lookup off its rungs to the
|
|
117
|
+
lowest available level; the session's bias-snapshot validator already rejects
|
|
118
|
+
such values before they reach the resolver. Coverage:
|
|
119
|
+
`packages/coding-agent/test/reasoning-router-resolver-contract.test.ts`.
|
|
120
|
+
|
|
47
121
|
## Working-tree shared run budgets
|
|
48
122
|
|
|
49
123
|
The SDK `prompt(..., { runBudget })` path now shares a monotonic deadline and
|
|
@@ -141,16 +215,19 @@ and places conflicting calls in later levels. Before a level executes, OMK autho
|
|
|
141
215
|
re-plans with post-hook arguments so a hook cannot silently invalidate the
|
|
142
216
|
original claim plan.
|
|
143
217
|
|
|
144
|
-
The live executor uses level barriers.
|
|
145
|
-
|
|
146
|
-
|
|
218
|
+
The live executor uses a ready frontier, not level barriers.
|
|
219
|
+
`assignDagDependencies()` computes the predecessor graph and the loop feeds it
|
|
220
|
+
to `runDagFrontier()`, which admits a node as soon as its predecessors settle
|
|
221
|
+
and a resident slot frees up. Each candidate is still authorized and
|
|
222
|
+
re-planned from post-hook arguments before execution.
|
|
147
223
|
|
|
148
224
|
Evidence:
|
|
149
225
|
|
|
150
226
|
- `packages/agent/src/tool-dag-scheduler.ts`: `assignDagLevels`,
|
|
151
227
|
`assignDagDependencies`, `scheduleDagLevels`
|
|
152
228
|
- `packages/agent/src/agent-loop.ts`: `executeToolCallsDagLevels`,
|
|
153
|
-
`
|
|
229
|
+
`runDagFrontier` (predecessor-settle ready queue; the former
|
|
230
|
+
`runDagLevelCalls` level barrier was replaced)
|
|
154
231
|
- `packages/agent/test/tool-dag-scheduler*.test.ts`
|
|
155
232
|
- `packages/agent/test/tool-dag-dependencies.test.ts`
|
|
156
233
|
|
|
@@ -174,21 +251,93 @@ The planner:
|
|
|
174
251
|
4. scores optional items for relevance, recency, evidence, redundancy,
|
|
175
252
|
priority, and full-text token cost;
|
|
176
253
|
5. sorts optional items by
|
|
177
|
-
`density -> effectiveScore -> priorityRank -> fullTokens -> id`;
|
|
254
|
+
`density -> effectiveScore -> priorityRank -> fullTokens -> id`;
|
|
178
255
|
6. selects a full, summary, headroom-compressed, pointer, or omitted
|
|
179
|
-
representation that fits
|
|
256
|
+
representation that fits (floor pass per tier, then the global pass);
|
|
257
|
+
7. lets each still-omitted item buy its cheapest admissible form by stepping
|
|
258
|
+
selected items down to theirs, applied all-or-nothing; and
|
|
259
|
+
8. re-offers each remaining item its costlier representations against the
|
|
260
|
+
leftover budget and promotes only when the selection policy prefers the
|
|
261
|
+
costlier form and it fits both the tier ceiling and the global remainder
|
|
262
|
+
(`context-budget-v2-global-pass.ts`; never touches hard items).
|
|
263
|
+
|
|
264
|
+
Breadth is settled before quality: step 7 runs before step 8, so an exchange
|
|
265
|
+
that admits another item cannot be undone by a promotion.
|
|
180
266
|
|
|
181
267
|
Density divides effective score by the cheapest non-omit representation
|
|
182
268
|
(`admissibleTokens`), not by full-text size. This avoids penalizing an item that
|
|
183
269
|
can be represented by a small evidence pointer. Stable item IDs and explicit
|
|
184
|
-
selection policy `sel-
|
|
270
|
+
selection policy `sel-3` make tie-breaking and plan-cache invalidation
|
|
185
271
|
deterministic.
|
|
186
272
|
|
|
273
|
+
### Representation cost accounting (2026-09-19 audit F01/F03)
|
|
274
|
+
|
|
275
|
+
Every derived representation is priced by counting the string it materializes
|
|
276
|
+
with the planner's own token counter, the same counter that priced the full
|
|
277
|
+
text. The former `ceil(0.15 * full) + 8` summary price was a compression target,
|
|
278
|
+
not a cost: a 100-character "summary" identical to its source was recorded at
|
|
279
|
+
12 tokens against the source's 25 and admitted into a 12-token budget. A
|
|
280
|
+
representation whose text equals the source, or whose counted cost is not below
|
|
281
|
+
the full text, is no longer offered. `createPlannedItems` stores the priced
|
|
282
|
+
candidates on each planned item so ranking and selection read one cost; the
|
|
283
|
+
selection policy token moved from `sel-2` to `sel-3` so plans and
|
|
284
|
+
representation entries cached under the old prices are not served.
|
|
285
|
+
|
|
286
|
+
### Representation exchange (2026-09-19 audit F04)
|
|
287
|
+
|
|
288
|
+
Items are ranked by their cheapest admissible representation but admitted at
|
|
289
|
+
whatever the selector prefers, so a high-priority item priced at 10 for ordering
|
|
290
|
+
could consume 100 and displace two 45-token items. On the audit's counterexample
|
|
291
|
+
that scored 121 against 169.5 for `A-pointer + B-full + C-full` at the same 100
|
|
292
|
+
tokens, measured with the selector's own preference score.
|
|
293
|
+
|
|
294
|
+
The repair is bounded. After the global pass, each still-omitted item in rank
|
|
295
|
+
order may pay for its admission by stepping already-selected items down to their
|
|
296
|
+
cheapest admissible representation. Donors yield in order of the preference the
|
|
297
|
+
policy loses per token freed, the whole exchange is applied or none of it is,
|
|
298
|
+
and a donor is skipped when stepping it down would drop its tier below the floor
|
|
299
|
+
it is currently honouring. An item admitted this way leaves the omitted list, so
|
|
300
|
+
it is never reported as both included and omitted.
|
|
301
|
+
|
|
302
|
+
This is not joint `(item, representation)` optimization: donor choice is greedy
|
|
303
|
+
and admission order is unchanged, so the planner can still trail a brute-force
|
|
304
|
+
oracle on instances the bounded repair cannot reach. It does reach the oracle on
|
|
305
|
+
the audit counterexample, and a 200-instance randomized property check holds
|
|
306
|
+
feasibility (global and tier caps), disjoint included/omitted sets, and
|
|
307
|
+
maximality: no omitted item still fits at its cheapest admissible form.
|
|
308
|
+
|
|
309
|
+
The quality-policy field `preferFullForHighPriority` is deprecated: nothing
|
|
310
|
+
reads it, and flipping it changes no candidate or choice
|
|
311
|
+
(`context-budget-quality-policy-semantics.test.ts`). The priority weight already
|
|
312
|
+
prefers full text for high-priority items.
|
|
313
|
+
|
|
314
|
+
Evidence:
|
|
315
|
+
|
|
316
|
+
- `packages/coding-agent/test/context-budget-representation-accounting.test.ts`
|
|
317
|
+
- `packages/coding-agent/test/context-budget-representation-exchange.test.ts`
|
|
318
|
+
- `packages/coding-agent/test/context-budget-representation-promotion.test.ts`
|
|
319
|
+
- `packages/coding-agent/test/context-budget-quality-policy-semantics.test.ts`
|
|
320
|
+
|
|
187
321
|
When enabled, representation and negative-result entries persist under
|
|
188
322
|
`.omk/cache/context-budget-v2`; plan entries remain session-memory-only.
|
|
189
323
|
`OMK_CONTEXT_GOVERNOR_CACHE=memory` keeps every cache entry in session memory,
|
|
190
324
|
and `OMK_CONTEXT_GOVERNOR_CACHE_DIR` relocates the representation snapshot.
|
|
191
325
|
|
|
326
|
+
Cache keys are bound to the counter that produced the prices. Two of the three
|
|
327
|
+
layers bind it implicitly — a plan key hashes each planned item's token counts,
|
|
328
|
+
and an exact representation key hashes a fingerprint containing
|
|
329
|
+
`estimatedTokens` — but the materialized (semantic) key is bucketed at 100
|
|
330
|
+
tokens and carries the tokenizer as a field, and no caller passed one, so every
|
|
331
|
+
counter shared the static `heuristic-v1` key space and the entry's
|
|
332
|
+
`tokenizer_mismatch` check compared that constant against itself. Since F01 made
|
|
333
|
+
every representation price counter-dependent, that let one estimator's price be
|
|
334
|
+
admitted in another estimator's run. The planner now resolves the key's
|
|
335
|
+
tokenizer from the counter's own `adapterId`
|
|
336
|
+
(`resolveEffectiveTokenizerIdV2`, probing a non-empty string because
|
|
337
|
+
`countText("")` short-circuits to the fallback estimator); an explicit
|
|
338
|
+
`tokenizerId` input still wins. Coverage:
|
|
339
|
+
`packages/coding-agent/test/context-budget-cache-tokenizer-binding.test.ts`.
|
|
340
|
+
|
|
192
341
|
Evidence:
|
|
193
342
|
|
|
194
343
|
- `packages/coding-agent/src/core/context-budget-v2-planner.ts`
|