open-multi-agent-kit 0.98.4 → 0.99.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +36 -0
- package/README.md +4 -2
- package/dist/cli/help.d.ts.map +1 -1
- package/dist/cli/help.js +1 -0
- package/dist/cli/help.js.map +1 -1
- package/dist/commands/run-command.d.ts +2 -1
- package/dist/commands/run-command.d.ts.map +1 -1
- package/dist/commands/run-command.js +9 -1
- package/dist/commands/run-command.js.map +1 -1
- package/dist/commands/verified-run-cli.d.ts +6 -0
- package/dist/commands/verified-run-cli.d.ts.map +1 -0
- package/dist/commands/verified-run-cli.js +190 -0
- package/dist/commands/verified-run-cli.js.map +1 -0
- package/dist/core/agent-session-services.d.ts +8 -1
- package/dist/core/agent-session-services.d.ts.map +1 -1
- package/dist/core/agent-session-services.js +41 -0
- package/dist/core/agent-session-services.js.map +1 -1
- package/dist/core/agent-session.d.ts +7 -8
- package/dist/core/agent-session.d.ts.map +1 -1
- package/dist/core/agent-session.js +80 -43
- package/dist/core/agent-session.js.map +1 -1
- package/dist/core/devin-harness-dispatch.d.ts +12 -0
- package/dist/core/devin-harness-dispatch.d.ts.map +1 -0
- package/dist/core/devin-harness-dispatch.js +12 -0
- package/dist/core/devin-harness-dispatch.js.map +1 -0
- package/dist/core/devin-harness.d.ts +53 -0
- package/dist/core/devin-harness.d.ts.map +1 -0
- package/dist/core/devin-harness.js +112 -0
- package/dist/core/devin-harness.js.map +1 -0
- package/dist/core/domain-dispatch.d.ts +4 -1
- package/dist/core/domain-dispatch.d.ts.map +1 -1
- package/dist/core/domain-dispatch.js +5 -0
- package/dist/core/domain-dispatch.js.map +1 -1
- package/dist/core/domain-loadouts-provider-harness.d.ts +12 -0
- package/dist/core/domain-loadouts-provider-harness.d.ts.map +1 -0
- package/dist/core/domain-loadouts-provider-harness.js +122 -0
- package/dist/core/domain-loadouts-provider-harness.js.map +1 -0
- package/dist/core/domain-loadouts.d.ts +2 -33
- package/dist/core/domain-loadouts.d.ts.map +1 -1
- package/dist/core/domain-loadouts.js +3 -56
- package/dist/core/domain-loadouts.js.map +1 -1
- package/dist/core/domain-profile.d.ts +40 -0
- package/dist/core/domain-profile.d.ts.map +1 -0
- package/dist/core/domain-profile.js +7 -0
- package/dist/core/domain-profile.js.map +1 -0
- package/dist/core/grok-harness-dispatch.d.ts +6 -20
- package/dist/core/grok-harness-dispatch.d.ts.map +1 -1
- package/dist/core/grok-harness-dispatch.js +6 -55
- package/dist/core/grok-harness-dispatch.js.map +1 -1
- package/dist/core/grok-harness.d.ts +7 -9
- package/dist/core/grok-harness.d.ts.map +1 -1
- package/dist/core/grok-harness.js +9 -23
- package/dist/core/grok-harness.js.map +1 -1
- package/dist/core/harness-skills.d.ts +20 -0
- package/dist/core/harness-skills.d.ts.map +1 -0
- package/dist/core/harness-skills.js +36 -0
- package/dist/core/harness-skills.js.map +1 -0
- package/dist/core/loadout-runtime-state.d.ts +25 -0
- package/dist/core/loadout-runtime-state.d.ts.map +1 -0
- package/dist/core/loadout-runtime-state.js +39 -0
- package/dist/core/loadout-runtime-state.js.map +1 -0
- package/dist/core/loadout-runtime.d.ts +2 -12
- package/dist/core/loadout-runtime.d.ts.map +1 -1
- package/dist/core/loadout-runtime.js +6 -13
- package/dist/core/loadout-runtime.js.map +1 -1
- package/dist/core/model-resolver.d.ts +1 -41
- package/dist/core/model-resolver.d.ts.map +1 -1
- package/dist/core/model-resolver.js +2 -49
- package/dist/core/model-resolver.js.map +1 -1
- package/dist/core/prompt-settlement.d.ts +4 -3
- package/dist/core/prompt-settlement.d.ts.map +1 -1
- package/dist/core/prompt-settlement.js.map +1 -1
- package/dist/core/provider-default-models.d.ts +42 -0
- package/dist/core/provider-default-models.d.ts.map +1 -0
- package/dist/core/provider-default-models.js +50 -0
- package/dist/core/provider-default-models.js.map +1 -0
- package/dist/core/provider-display-names.d.ts.map +1 -1
- package/dist/core/provider-display-names.js +1 -0
- package/dist/core/provider-display-names.js.map +1 -1
- package/dist/core/provider-harness-dispatch.d.ts +63 -0
- package/dist/core/provider-harness-dispatch.d.ts.map +1 -0
- package/dist/core/provider-harness-dispatch.js +60 -0
- package/dist/core/provider-harness-dispatch.js.map +1 -0
- package/dist/core/provider-usage-devin.d.ts +16 -0
- package/dist/core/provider-usage-devin.d.ts.map +1 -0
- package/dist/core/provider-usage-devin.js +61 -0
- package/dist/core/provider-usage-devin.js.map +1 -0
- package/dist/core/provider-usage-text.d.ts +5 -0
- package/dist/core/provider-usage-text.d.ts.map +1 -0
- package/dist/core/provider-usage-text.js +15 -0
- package/dist/core/provider-usage-text.js.map +1 -0
- package/dist/core/provider-usage-types.d.ts +2 -1
- package/dist/core/provider-usage-types.d.ts.map +1 -1
- package/dist/core/provider-usage-types.js.map +1 -1
- package/dist/core/provider-usage.d.ts +1 -2
- package/dist/core/provider-usage.d.ts.map +1 -1
- package/dist/core/provider-usage.js +12 -7
- package/dist/core/provider-usage.js.map +1 -1
- package/dist/core/run-budget-policy.d.ts +18 -0
- package/dist/core/run-budget-policy.d.ts.map +1 -0
- package/dist/core/run-budget-policy.js +52 -0
- package/dist/core/run-budget-policy.js.map +1 -0
- package/dist/core/run-budget.d.ts +33 -0
- package/dist/core/run-budget.d.ts.map +1 -0
- package/dist/core/run-budget.js +89 -0
- package/dist/core/run-budget.js.map +1 -0
- package/dist/core/run-execution-api.d.ts +15 -0
- package/dist/core/run-execution-api.d.ts.map +1 -0
- package/dist/core/run-execution-api.js +7 -0
- package/dist/core/run-execution-api.js.map +1 -0
- package/dist/core/run-journal.d.ts.map +1 -1
- package/dist/core/run-journal.js +30 -208
- package/dist/core/run-journal.js.map +1 -1
- package/dist/core/sdk.d.ts.map +1 -1
- package/dist/core/sdk.js +23 -18
- package/dist/core/sdk.js.map +1 -1
- package/dist/core/session-bash-service.d.ts +2 -2
- package/dist/core/session-bash-service.d.ts.map +1 -1
- package/dist/core/session-bash-service.js +21 -9
- package/dist/core/session-bash-service.js.map +1 -1
- package/dist/core/session-failure-cause.d.ts.map +1 -1
- package/dist/core/session-failure-cause.js +5 -0
- package/dist/core/session-failure-cause.js.map +1 -1
- package/dist/core/session-prompt-lifecycle.d.ts +25 -0
- package/dist/core/session-prompt-lifecycle.d.ts.map +1 -0
- package/dist/core/session-prompt-lifecycle.js +87 -0
- package/dist/core/session-prompt-lifecycle.js.map +1 -0
- package/dist/core/session-run-budget.d.ts +28 -0
- package/dist/core/session-run-budget.d.ts.map +1 -0
- package/dist/core/session-run-budget.js +127 -0
- package/dist/core/session-run-budget.js.map +1 -0
- package/dist/core/session-run-termination.d.ts.map +1 -1
- package/dist/core/session-run-termination.js +14 -3
- package/dist/core/session-run-termination.js.map +1 -1
- package/dist/core/session-termination-types.d.ts +97 -0
- package/dist/core/session-termination-types.d.ts.map +1 -0
- package/dist/core/session-termination-types.js +26 -0
- package/dist/core/session-termination-types.js.map +1 -0
- package/dist/core/session-termination.d.ts +3 -97
- package/dist/core/session-termination.d.ts.map +1 -1
- package/dist/core/session-termination.js +24 -28
- package/dist/core/session-termination.js.map +1 -1
- package/dist/core/slash-commands.d.ts.map +1 -1
- package/dist/core/slash-commands.js +1 -0
- package/dist/core/slash-commands.js.map +1 -1
- package/dist/core/subagent-lane-launcher.d.ts +3 -2
- package/dist/core/subagent-lane-launcher.d.ts.map +1 -1
- package/dist/core/subagent-lane-launcher.js +24 -13
- package/dist/core/subagent-lane-launcher.js.map +1 -1
- package/dist/core/verified-run/broker.d.ts +26 -0
- package/dist/core/verified-run/broker.d.ts.map +1 -0
- package/dist/core/verified-run/broker.js +210 -0
- package/dist/core/verified-run/broker.js.map +1 -0
- package/dist/core/verified-run/candidate.d.ts +24 -0
- package/dist/core/verified-run/candidate.d.ts.map +1 -0
- package/dist/core/verified-run/candidate.js +167 -0
- package/dist/core/verified-run/candidate.js.map +1 -0
- package/dist/core/verified-run/check-receipt.d.ts +18 -0
- package/dist/core/verified-run/check-receipt.d.ts.map +1 -0
- package/dist/core/verified-run/check-receipt.js +92 -0
- package/dist/core/verified-run/check-receipt.js.map +1 -0
- package/dist/core/verified-run/coordinator.d.ts +39 -0
- package/dist/core/verified-run/coordinator.d.ts.map +1 -0
- package/dist/core/verified-run/coordinator.js +190 -0
- package/dist/core/verified-run/coordinator.js.map +1 -0
- package/dist/core/verified-run/dag-candidates.d.ts +8 -0
- package/dist/core/verified-run/dag-candidates.d.ts.map +1 -0
- package/dist/core/verified-run/dag-candidates.js +65 -0
- package/dist/core/verified-run/dag-candidates.js.map +1 -0
- package/dist/core/verified-run/dag-phase.d.ts +7 -0
- package/dist/core/verified-run/dag-phase.d.ts.map +1 -0
- package/dist/core/verified-run/dag-phase.js +130 -0
- package/dist/core/verified-run/dag-phase.js.map +1 -0
- package/dist/core/verified-run/dag-projection.d.ts +7 -0
- package/dist/core/verified-run/dag-projection.d.ts.map +1 -0
- package/dist/core/verified-run/dag-projection.js +103 -0
- package/dist/core/verified-run/dag-projection.js.map +1 -0
- package/dist/core/verified-run/dag-recovery.d.ts +13 -0
- package/dist/core/verified-run/dag-recovery.d.ts.map +1 -0
- package/dist/core/verified-run/dag-recovery.js +98 -0
- package/dist/core/verified-run/dag-recovery.js.map +1 -0
- package/dist/core/verified-run/dag-retry-projection.d.ts +9 -0
- package/dist/core/verified-run/dag-retry-projection.d.ts.map +1 -0
- package/dist/core/verified-run/dag-retry-projection.js +43 -0
- package/dist/core/verified-run/dag-retry-projection.js.map +1 -0
- package/dist/core/verified-run/dag-types.d.ts +58 -0
- package/dist/core/verified-run/dag-types.d.ts.map +1 -0
- package/dist/core/verified-run/dag-types.js +2 -0
- package/dist/core/verified-run/dag-types.js.map +1 -0
- package/dist/core/verified-run/event-parser.d.ts +3 -0
- package/dist/core/verified-run/event-parser.d.ts.map +1 -0
- package/dist/core/verified-run/event-parser.js +154 -0
- package/dist/core/verified-run/event-parser.js.map +1 -0
- package/dist/core/verified-run/events.d.ts +4 -0
- package/dist/core/verified-run/events.d.ts.map +1 -0
- package/dist/core/verified-run/events.js +3 -0
- package/dist/core/verified-run/events.js.map +1 -0
- package/dist/core/verified-run/evidence-binding.d.ts +18 -0
- package/dist/core/verified-run/evidence-binding.d.ts.map +1 -0
- package/dist/core/verified-run/evidence-binding.js +81 -0
- package/dist/core/verified-run/evidence-binding.js.map +1 -0
- package/dist/core/verified-run/evidence.d.ts +27 -0
- package/dist/core/verified-run/evidence.d.ts.map +1 -0
- package/dist/core/verified-run/evidence.js +100 -0
- package/dist/core/verified-run/evidence.js.map +1 -0
- package/dist/core/verified-run/journal.d.ts +32 -0
- package/dist/core/verified-run/journal.d.ts.map +1 -0
- package/dist/core/verified-run/journal.js +103 -0
- package/dist/core/verified-run/journal.js.map +1 -0
- package/dist/core/verified-run/namespace-identity.d.ts +11 -0
- package/dist/core/verified-run/namespace-identity.d.ts.map +1 -0
- package/dist/core/verified-run/namespace-identity.js +77 -0
- package/dist/core/verified-run/namespace-identity.js.map +1 -0
- package/dist/core/verified-run/owned-execution.d.ts +22 -0
- package/dist/core/verified-run/owned-execution.d.ts.map +1 -0
- package/dist/core/verified-run/owned-execution.js +49 -0
- package/dist/core/verified-run/owned-execution.js.map +1 -0
- package/dist/core/verified-run/phase-context.d.ts +9 -0
- package/dist/core/verified-run/phase-context.d.ts.map +1 -0
- package/dist/core/verified-run/phase-context.js +2 -0
- package/dist/core/verified-run/phase-context.js.map +1 -0
- package/dist/core/verified-run/process-gate.d.ts +5 -0
- package/dist/core/verified-run/process-gate.d.ts.map +1 -0
- package/dist/core/verified-run/process-gate.js +31 -0
- package/dist/core/verified-run/process-gate.js.map +1 -0
- package/dist/core/verified-run/process-projection.d.ts +13 -0
- package/dist/core/verified-run/process-projection.d.ts.map +1 -0
- package/dist/core/verified-run/process-projection.js +82 -0
- package/dist/core/verified-run/process-projection.js.map +1 -0
- package/dist/core/verified-run/projection.d.ts +4 -0
- package/dist/core/verified-run/projection.d.ts.map +1 -0
- package/dist/core/verified-run/projection.js +175 -0
- package/dist/core/verified-run/projection.js.map +1 -0
- package/dist/core/verified-run/recovery-clock.d.ts +19 -0
- package/dist/core/verified-run/recovery-clock.d.ts.map +1 -0
- package/dist/core/verified-run/recovery-clock.js +69 -0
- package/dist/core/verified-run/recovery-clock.js.map +1 -0
- package/dist/core/verified-run/recovery-command.d.ts +11 -0
- package/dist/core/verified-run/recovery-command.d.ts.map +1 -0
- package/dist/core/verified-run/recovery-command.js +73 -0
- package/dist/core/verified-run/recovery-command.js.map +1 -0
- package/dist/core/verified-run/recovery-projection.d.ts +12 -0
- package/dist/core/verified-run/recovery-projection.d.ts.map +1 -0
- package/dist/core/verified-run/recovery-projection.js +88 -0
- package/dist/core/verified-run/recovery-projection.js.map +1 -0
- package/dist/core/verified-run/recovery.d.ts +12 -0
- package/dist/core/verified-run/recovery.d.ts.map +1 -0
- package/dist/core/verified-run/recovery.js +115 -0
- package/dist/core/verified-run/recovery.js.map +1 -0
- package/dist/core/verified-run/run-types.d.ts +104 -0
- package/dist/core/verified-run/run-types.d.ts.map +1 -0
- package/dist/core/verified-run/run-types.js +2 -0
- package/dist/core/verified-run/run-types.js.map +1 -0
- package/dist/core/verified-run/scripted-writer.d.ts +15 -0
- package/dist/core/verified-run/scripted-writer.d.ts.map +1 -0
- package/dist/core/verified-run/scripted-writer.js +94 -0
- package/dist/core/verified-run/scripted-writer.js.map +1 -0
- package/dist/core/verified-run/session-port.d.ts +27 -0
- package/dist/core/verified-run/session-port.d.ts.map +1 -0
- package/dist/core/verified-run/session-port.js +2 -0
- package/dist/core/verified-run/session-port.js.map +1 -0
- package/dist/core/verified-run/storage.d.ts +15 -0
- package/dist/core/verified-run/storage.d.ts.map +1 -0
- package/dist/core/verified-run/storage.js +96 -0
- package/dist/core/verified-run/storage.js.map +1 -0
- package/dist/core/verified-run/task-execution.d.ts +9 -0
- package/dist/core/verified-run/task-execution.d.ts.map +1 -0
- package/dist/core/verified-run/task-execution.js +29 -0
- package/dist/core/verified-run/task-execution.js.map +1 -0
- package/dist/core/verified-run/verification-phase.d.ts +5 -0
- package/dist/core/verified-run/verification-phase.d.ts.map +1 -0
- package/dist/core/verified-run/verification-phase.js +56 -0
- package/dist/core/verified-run/verification-phase.js.map +1 -0
- package/dist/core/verified-run/work-recovery.d.ts +8 -0
- package/dist/core/verified-run/work-recovery.d.ts.map +1 -0
- package/dist/core/verified-run/work-recovery.js +44 -0
- package/dist/core/verified-run/work-recovery.js.map +1 -0
- package/dist/core/verified-run/writer-completion.d.ts +4 -0
- package/dist/core/verified-run/writer-completion.d.ts.map +1 -0
- package/dist/core/verified-run/writer-completion.js +27 -0
- package/dist/core/verified-run/writer-completion.js.map +1 -0
- package/dist/core/verified-run/writer-phase.d.ts +8 -0
- package/dist/core/verified-run/writer-phase.d.ts.map +1 -0
- package/dist/core/verified-run/writer-phase.js +76 -0
- package/dist/core/verified-run/writer-phase.js.map +1 -0
- package/dist/core/verified-run/writer-projection.d.ts +11 -0
- package/dist/core/verified-run/writer-projection.d.ts.map +1 -0
- package/dist/core/verified-run/writer-projection.js +71 -0
- package/dist/core/verified-run/writer-projection.js.map +1 -0
- package/dist/core/verified-run/writer-recovery.d.ts +16 -0
- package/dist/core/verified-run/writer-recovery.d.ts.map +1 -0
- package/dist/core/verified-run/writer-recovery.js +107 -0
- package/dist/core/verified-run/writer-recovery.js.map +1 -0
- package/dist/core/workload-permit-pool.d.ts +1 -3
- package/dist/core/workload-permit-pool.d.ts.map +1 -1
- package/dist/core/workload-permit-pool.js +11 -11
- package/dist/core/workload-permit-pool.js.map +1 -1
- package/dist/index.d.ts +3 -2
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +2 -2
- package/dist/index.js.map +1 -1
- package/dist/main.d.ts.map +1 -1
- package/dist/main.js +2 -8
- package/dist/main.js.map +1 -1
- package/dist/modes/interactive/components/session-failure.d.ts +13 -0
- package/dist/modes/interactive/components/session-failure.d.ts.map +1 -0
- package/dist/modes/interactive/components/session-failure.js +81 -0
- package/dist/modes/interactive/components/session-failure.js.map +1 -0
- package/dist/modes/interactive/interactive-mode.d.ts +1 -0
- package/dist/modes/interactive/interactive-mode.d.ts.map +1 -1
- package/dist/modes/interactive/interactive-mode.js +47 -36
- package/dist/modes/interactive/interactive-mode.js.map +1 -1
- package/dist/modes/interactive/resource-description.d.ts +3 -0
- package/dist/modes/interactive/resource-description.d.ts.map +1 -0
- package/dist/modes/interactive/resource-description.js +15 -0
- package/dist/modes/interactive/resource-description.js.map +1 -0
- package/dist/modes/interactive/tui-diagnostics.d.ts +36 -0
- package/dist/modes/interactive/tui-diagnostics.d.ts.map +1 -0
- package/dist/modes/interactive/tui-diagnostics.js +105 -0
- package/dist/modes/interactive/tui-diagnostics.js.map +1 -0
- package/dist/modes/interactive/tui-runtime-info.d.ts +20 -0
- package/dist/modes/interactive/tui-runtime-info.d.ts.map +1 -0
- package/dist/modes/interactive/tui-runtime-info.js +76 -0
- package/dist/modes/interactive/tui-runtime-info.js.map +1 -0
- package/docs/development.md +10 -0
- package/docs/devin-harness.md +121 -0
- package/docs/docs.json +4 -0
- package/docs/environment-variables.md +2 -1
- package/docs/index.md +1 -0
- package/docs/keybindings.md +1 -1
- package/docs/loadout-domains/README.md +2 -1
- package/docs/loadout-domains/devin-harness.md +72 -0
- package/docs/metrics.md +57 -16
- package/docs/providers.md +76 -0
- package/docs/release-audit-0.98.5.md +101 -0
- package/docs/release-audit-0.99.0.md +68 -0
- package/docs/run-protocol.md +34 -1
- package/docs/runtime-algorithms.md +30 -1
- package/docs/sdk.md +198 -1
- package/docs/settings.md +1 -1
- package/docs/startup-resource-labels-testing.md +47 -0
- package/docs/tb21-audit.md +18 -4
- package/docs/usage.md +59 -0
- package/docs/verified-run-remaining-design.md +881 -0
- package/docs/verified-run-testing.md +508 -0
- package/docs/verified-run.md +409 -0
- package/examples/extensions/custom-provider-anthropic/package-lock.json +2 -2
- package/examples/extensions/custom-provider-anthropic/package.json +1 -1
- package/examples/extensions/custom-provider-gitlab-duo/package.json +1 -1
- package/examples/extensions/gondolin/package-lock.json +2 -2
- package/examples/extensions/gondolin/package.json +1 -1
- package/examples/extensions/sandbox/package-lock.json +2 -2
- package/examples/extensions/sandbox/package.json +1 -1
- package/examples/extensions/with-deps/package-lock.json +2 -2
- package/examples/extensions/with-deps/package.json +1 -1
- package/npm-shrinkwrap.json +18 -18
- package/package.json +6 -6
package/docs/metrics.md
CHANGED
|
@@ -160,7 +160,7 @@ node scripts/tb-mini-suite.mjs --json # feed a runner
|
|
|
160
160
|
node scripts/tb-mini-suite.mjs --seed 7 # a different fixed subset
|
|
161
161
|
```
|
|
162
162
|
|
|
163
|
-
With identical task metadata, seed, size, and collation, selection is repeatable.
|
|
163
|
+
With identical task metadata, selection version, seed, size, and collation, selection is repeatable.
|
|
164
164
|
The default 15-task subset oversamples easy tasks and prioritizes shorter expert
|
|
165
165
|
time estimates; it is a regression signal, not a population-representative score
|
|
166
166
|
or an agent runtime bound. Selection alone is not a capability result. Scoring
|
|
@@ -171,12 +171,35 @@ population. Missing difficulty quotas are filled from unselected tasks using the
|
|
|
171
171
|
same ordering, so a valid request returns exactly that many distinct tasks.
|
|
172
172
|
`--seed` accepts integers from `0` through `4294967295`. Invalid or missing option
|
|
173
173
|
values and oversized requests exit with code `2`; absent, empty, or non-directory
|
|
174
|
-
task paths exit with code `1`.
|
|
174
|
+
task paths exit with code `1`.
|
|
175
|
+
|
|
176
|
+
The JSON output now declares `selectionVersion: 2`. Missing, empty, nonfinite, or
|
|
177
|
+
negative expert-time estimates are `null`, not zero. Within each difficulty band
|
|
178
|
+
and in quota refill, known estimates sort before unknown estimates. A genuine zero
|
|
179
|
+
or fractional estimate remains valid. Difficulty quotas still take precedence, so
|
|
180
|
+
unknown-estimate tasks can be selected to fill a band.
|
|
181
|
+
|
|
182
|
+
`knownExpertMinutes` sums known estimates; `unknownExpertEstimates` counts selected
|
|
183
|
+
tasks with unknown estimates. `totalExpertMinutes` is `null` when any selected
|
|
184
|
+
estimate is unknown. A nonfinite sum is refused with exit `1`, not serialized as
|
|
185
|
+
an apparently missing total. Expert estimates are not agent timeout limits.
|
|
186
|
+
|
|
187
|
+
For sizes 1–2, available slots go first to the highest-weight difficulty bands
|
|
188
|
+
(medium, then hard), with normal refill if those bands are unavailable. Task names
|
|
189
|
+
no longer decide which excess band quota is discarded. The normal-size quota rule
|
|
190
|
+
is preserved. Human-readable counts use a `Map`, including for labels such as
|
|
191
|
+
`__proto__` that overlap JavaScript object properties.
|
|
192
|
+
|
|
193
|
+
This is a selection and output-contract change: default membership can change when
|
|
194
|
+
metadata is incomplete, and the prior default JSON digest is historical only.
|
|
195
|
+
Freeze new task manifests before comparison; do not combine version 1 and 2 runs
|
|
196
|
+
as if their selection policy were identical. The limited flat TOML field reader
|
|
197
|
+
is unchanged; this does not add general TOML syntax support.
|
|
175
198
|
|
|
176
199
|
Run the offline CLI regression tests without downloading tasks or calling models:
|
|
177
200
|
|
|
178
201
|
```bash
|
|
179
|
-
node --test scripts/test/tb-mini-suite.test.mjs
|
|
202
|
+
node --test --test-concurrency=1 scripts/test/tb-mini-suite.test.mjs scripts/test/tb-mini-suite-ranking.test.mjs
|
|
180
203
|
```
|
|
181
204
|
|
|
182
205
|
See [the harness roadmap](../../../ROADMAP.md) for the dated OMK versus Terminus-2
|
|
@@ -187,8 +210,8 @@ criteria. Planned runtime improvements are not measured benchmark gains.
|
|
|
187
210
|
|
|
188
211
|
The checkout-only `scripts/tb21-audit.mjs` audits explicitly selected Harbor jobs
|
|
189
212
|
against a caller-pinned manifest digest. It rejects duplicate tasks/trials, missing
|
|
190
|
-
results or costs, mismatched task checksums or configured model labels,
|
|
191
|
-
contradictory success records. It never starts a model, picks the latest job, joins
|
|
213
|
+
results or costs, mismatched task checksums or configured model labels,
|
|
214
|
+
missing/invalid completion times, and contradictory success records. It never starts a model, picks the latest job, joins
|
|
192
215
|
requests by timestamp, or rewrites evidence. See [TB 2.1 offline audit](tb21-audit.md)
|
|
193
216
|
for the schema, invocation, error codes, and limitations.
|
|
194
217
|
|
|
@@ -196,14 +219,32 @@ A complete audit means recorded outcomes passed these checks, not that every
|
|
|
196
219
|
provider request obeyed a single-model contract. Wire provenance, actual billing,
|
|
197
220
|
repeated-trial analysis, and statistical superiority need separate evidence.
|
|
198
221
|
|
|
199
|
-
###
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
222
|
+
### Output-limit validation: availability history
|
|
223
|
+
|
|
224
|
+
**2026-09-13 source snapshot (`ca75f4e5cc`):** the logical contract, CLI/SDK
|
|
225
|
+
wiring, and final Chat Completions model/output-limit checks are in the committed
|
|
226
|
+
source. [Model dispatch contracts](model-contract.md) defines their coverage;
|
|
227
|
+
[Verified Run](verified-run.md) describes the separate protected execution path.
|
|
228
|
+
Neither is a universal provider billing cap or a new controlled benchmark result.
|
|
229
|
+
See [the roadmap, section 16](../../../ROADMAP.md) for current local release-preparation checks.
|
|
230
|
+
|
|
231
|
+
The following paragraphs describe the earlier checkout only. Its missing modules
|
|
232
|
+
and failing collection are historical observations, not active release blockers.
|
|
233
|
+
|
|
234
|
+
**2026-09-08 follow-up:** the current worktree restores the logical contract and
|
|
235
|
+
connects it through the CLI/SDK, including SDK-stream summaries. See
|
|
236
|
+
[Model dispatch contracts](model-contract.md) and ROADMAP §13 for fresh evidence
|
|
237
|
+
and the remaining final-wire/accounting gaps. The warning below records the
|
|
238
|
+
preceding checkout, not the current availability of the restored module.
|
|
239
|
+
|
|
240
|
+
The previous worktree checkpoint tested positive-safe-integer validation of
|
|
241
|
+
`modelContract.maxOutputTokens` and explicit request `maxTokens`. During the
|
|
242
|
+
2026-09-08 re-verification, the checkout changed: `run-model-contract.ts` and the
|
|
243
|
+
corresponding `AgentLoopConfig.modelContract` surface were absent. The remaining
|
|
244
|
+
`model-contract-output-limit.test.ts` fails collection against that checkout.
|
|
245
|
+
|
|
246
|
+
Do not treat the historical passing tests as proof that this guard is currently
|
|
247
|
+
available. Restoring or porting the runtime contract requires an explicit source
|
|
248
|
+
baseline decision and fresh send-boundary tests. No missing code was silently
|
|
249
|
+
recreated and no failing test was deleted. Full run-wide enforcement, including
|
|
250
|
+
omitted limits and compaction, remains unverified; see ROADMAP sections 11–12.
|
package/docs/providers.md
CHANGED
|
@@ -19,6 +19,7 @@ Use `/login` in interactive mode, then select a provider:
|
|
|
19
19
|
- Claude Pro/Max
|
|
20
20
|
- GitHub Copilot
|
|
21
21
|
- xAI Grok subscription OAuth
|
|
22
|
+
- Devin CLI subscription (SWE-2)
|
|
22
23
|
|
|
23
24
|
Run `/login` and choose a configured subscription provider to open its account picker. Select an existing account by its ChatGPT, Claude, or Google email when available, or choose **Add another account** to sign in with a new one. OMK keeps and refreshes each account independently, pins the provider to the account you select, and does not silently fail over to another subscription. `/model` remains dedicated to model selection.
|
|
24
25
|
|
|
@@ -67,6 +68,80 @@ and [plan endpoint separation](https://www.alibabacloud.com/help/en/model-studio
|
|
|
67
68
|
consulted 2026-09-08. The local tests exercise serialization, not account availability,
|
|
68
69
|
provider compliance, billing, or benchmark performance.
|
|
69
70
|
|
|
71
|
+
### Devin CLI
|
|
72
|
+
|
|
73
|
+
Run `/login devin`, select `devin/swe-2`, and use `/think medium`, `/think high`,
|
|
74
|
+
or `/think max`. For a new session after login:
|
|
75
|
+
|
|
76
|
+
```bash
|
|
77
|
+
omk --provider devin --model swe-2 --thinking max
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
The default is `medium`. Thinking-off and other levels are unsupported. `max`
|
|
81
|
+
is a reasoning level, not a requirement to buy the Devin Max subscription tier.
|
|
82
|
+
Your account's model access and quota still apply. For presets, effort guidance,
|
|
83
|
+
the 1M-token context budget, and the `devin-harness` loadout, see the
|
|
84
|
+
[Devin SWE-2 harness](devin-harness.md).
|
|
85
|
+
|
|
86
|
+
**Authentication.** OMK uses the CLI's PKCE flow at
|
|
87
|
+
`app.devin.ai/auth/cli/continue` and `api.devin.ai/auth/cli/token`. Its callback
|
|
88
|
+
binds only to `127.0.0.1:59653`, validates state, and closes after success,
|
|
89
|
+
rejection, cancellation, or a five-minute deadline. If the port is occupied,
|
|
90
|
+
paste the complete callback URL into OMK's local prompt. For remote sessions,
|
|
91
|
+
forward this loopback port to the machine running OMK.
|
|
92
|
+
|
|
93
|
+
Credentials use the existing protected `auth.json` storage and account picker.
|
|
94
|
+
Expired sessions require `/login devin` again: no refresh endpoint is verified,
|
|
95
|
+
and OMK does not renew the expiry locally. Use `/logout` to remove the stored
|
|
96
|
+
OMK credential. Alternatively, `DEVIN_API_KEY` accepts an already-owned CLI
|
|
97
|
+
session token, **not** a `cog_` REST API key. OMK does not install or invoke the
|
|
98
|
+
Devin agent, read browser cookies, or import another application's credentials.
|
|
99
|
+
|
|
100
|
+
**Transport and limits.** The Node-only `devin-agent` adapter uses Connect/protobuf
|
|
101
|
+
at the fixed HTTPS origin `https://server.codeium.com`. It exchanges the session
|
|
102
|
+
token for a user JWT and reads `GetCliModelConfigs` before each turn. The server's
|
|
103
|
+
SWE-2 family metadata supplies the effort's exact wire UID; OMK never invents a
|
|
104
|
+
`swe-2-max` ID or downgrades an unavailable route. Disabled, internal, ambiguous,
|
|
105
|
+
and fast-lane entries are excluded. The family's separate 1M-context entries form
|
|
106
|
+
a second lane that is selected only when the model's local `contextWindow` is
|
|
107
|
+
1,000,000 or more (the bundled default); a smaller budget uses the standard lane.
|
|
108
|
+
|
|
109
|
+
Text, thinking, tool calls/results, and reported usage enter the normal OMK
|
|
110
|
+
agent loop. Image input and enterprise-origin overrides are unsupported.
|
|
111
|
+
Requests reject redirects. Malformed, oversized, or unterminated streams fail;
|
|
112
|
+
remote error bodies are not copied into diagnostics. SDK `onPayload` receives
|
|
113
|
+
protobuf bytes without credential metadata.
|
|
114
|
+
|
|
115
|
+
The bundled 1,000,000-token context budget and 16,384-token output cap are
|
|
116
|
+
local defaults, **not published SWE-2 limits**. Output is further capped against
|
|
117
|
+
the authenticated catalog. If the selected lane declares a smaller context window,
|
|
118
|
+
the request fails and names the window; lower `contextWindow` through
|
|
119
|
+
`modelOverrides` in [models.json](models.md#per-model-overrides) rather than
|
|
120
|
+
expecting a silent downgrade. Zero catalog pricing means unpriced subscription
|
|
121
|
+
usage, not free inference. Quota percentages are not inferred.
|
|
122
|
+
|
|
123
|
+
**Account quota.** With a Devin credential configured, the status rail's USAGE
|
|
124
|
+
section calls `SeatManagementService/GetUserStatus` (session-token metadata, no
|
|
125
|
+
user JWT) and renders the plan's daily and weekly quota meters with reset times.
|
|
126
|
+
Accounts whose plan reports no quota windows show the plan name and credit
|
|
127
|
+
balances instead. The rail mirrors the CLI's `/usage` surface; it never sends
|
|
128
|
+
the session token anywhere except `server.codeium.com`.
|
|
129
|
+
|
|
130
|
+
**Verification.** Local tests exercise the public stream API, real loopback
|
|
131
|
+
callbacks with mocked token exchange, model selection, and protocol fixtures.
|
|
132
|
+
Live login, catalog compatibility, SWE-2 inference, and billing remain unverified
|
|
133
|
+
without a Devin account. Unauthenticated catalog probes returned HTTP 400.
|
|
134
|
+
Live tests require `DEVIN_API_KEY` and `LIVE_E2E=1` and consume subscription quota.
|
|
135
|
+
|
|
136
|
+
Sources consulted 2026-09-12 and 2026-09-13: the [SWE-2 announcement](https://cognition.com/blog/swe-2)
|
|
137
|
+
(2026-09-10; CLI availability and medium/high/max; no published context window), [CLI commands](https://docs.devin.ai/cli/reference/commands),
|
|
138
|
+
and the [official manifest](https://static.devin.ai/cli/current/manifest.json)
|
|
139
|
+
(observed identity `3000.10.21`). Protocol fields follow the third-party
|
|
140
|
+
[oh-my-pi snapshot](https://github.com/can1357/oh-my-pi/blob/942383f768c5f2f6a620fcab57326c0f59df623a/packages/catalog/src/discovery/devin-proto.ts),
|
|
141
|
+
not a stable public inference contract; attribution is in `packages/ai/DEVIN-NOTICE`.
|
|
142
|
+
Regenerate only this logical model, preserving every other catalog entry, with
|
|
143
|
+
`npm --prefix packages/ai run generate-models -- --devin-only`.
|
|
144
|
+
|
|
70
145
|
### OpenAI Codex
|
|
71
146
|
|
|
72
147
|
- Requires ChatGPT Plus or Pro subscription
|
|
@@ -111,6 +186,7 @@ omk
|
|
|
111
186
|
| Azure OpenAI Responses | `AZURE_OPENAI_API_KEY` | `azure-openai-responses` |
|
|
112
187
|
| OpenAI | `OPENAI_API_KEY` | `openai` |
|
|
113
188
|
| DeepSeek | `DEEPSEEK_API_KEY` | `deepseek` |
|
|
189
|
+
| Devin CLI session token | `DEVIN_API_KEY` | `devin` |
|
|
114
190
|
| NVIDIA NIM | `NVIDIA_API_KEY` | `nvidia` |
|
|
115
191
|
| Google Gemini | `GEMINI_API_KEY` | `google` |
|
|
116
192
|
| Mistral | `MISTRAL_API_KEY` | `mistral` |
|
|
@@ -0,0 +1,101 @@
|
|
|
1
|
+
# Release audit: v0.98.5
|
|
2
|
+
|
|
3
|
+
Date: 2026-09-12. This records candidate preparation and verification, not proof of
|
|
4
|
+
semantic correctness, comparative performance or completed publication.
|
|
5
|
+
|
|
6
|
+
## Scope and changelog audit
|
|
7
|
+
|
|
8
|
+
The candidate retains the public v0.98.4 ancestry and the committed implementation
|
|
9
|
+
range through `b84d0d9b8e`. The audit covers index-preserving Git checks, execution
|
|
10
|
+
ownership/shared budgets, protected verified runs and candidate recovery, input
|
|
11
|
+
checkpoint writer restart, static DAG retry, and TUI diagnostics. Tool timeout
|
|
12
|
+
messages now distinguish requested cancellation from observed termination and use
|
|
13
|
+
a monotonic teardown grace period.
|
|
14
|
+
|
|
15
|
+
Preparation adds release metadata, this audit, and focused verification fixes:
|
|
16
|
+
initial failure-card expansion is applied at construction, execution wrappers
|
|
17
|
+
preserve lazy context-sensitive timeouts, and the commit hook handles changed-file
|
|
18
|
+
status under Husky's errexit mode. Existing dirty provider, retry,
|
|
19
|
+
parallel-frontier, resource-label and TB changes are excluded.
|
|
20
|
+
Their working files are preserved, and their pending changelog entries stay out of
|
|
21
|
+
the candidate. Published v0.98.4 and earlier changelog bodies remain unchanged.
|
|
22
|
+
Model catalogs are reused rather than fetched or regenerated. Credentials, private
|
|
23
|
+
agent-home configuration and the user's running TUI are not changed.
|
|
24
|
+
|
|
25
|
+
## Version and artifact contract
|
|
26
|
+
|
|
27
|
+
All seven public packages target 0.98.5: `open-multi-agent-kit`, `omk-ai`,
|
|
28
|
+
`omk-agent-core`, `omk-tui`, `omk-protocol`, `omk-adaptorch-wpl` and
|
|
29
|
+
`omk-book-to-skill`. Root/example manifests and locks, internal dependency ranges,
|
|
30
|
+
the book compiler's source version constant, CLI shrinkwrap and README pointers
|
|
31
|
+
are synchronized. External dependency versions and integrity values are unchanged.
|
|
32
|
+
|
|
33
|
+
The release candidate is selected by path/hunk and inspected as a complete staged
|
|
34
|
+
diff. Verification uses an isolated candidate, not the mixed working tree. Temporary
|
|
35
|
+
homes contain no provider credentials; live inference and benchmarks are not run.
|
|
36
|
+
|
|
37
|
+
## Observed verification
|
|
38
|
+
|
|
39
|
+
- The diagnostics implementation passed type checking and 104 focused tests in its
|
|
40
|
+
staged snapshot. Its canonical installed launcher passed offline 120/80-column
|
|
41
|
+
inspection, save, failure, reload and quit checks without model prompts.
|
|
42
|
+
- Version preparation passed 13-manifest/root-lock parity, source-version tests,
|
|
43
|
+
book metadata tests and the canonical `omk --version` check for 0.98.5.
|
|
44
|
+
- The isolated candidate build and `npm run check` exited 0. Module-size and
|
|
45
|
+
import-cycle baselines were not raised.
|
|
46
|
+
- The final full offline run exited 0: 8,181 passed, 837 environment/live-condition
|
|
47
|
+
skips, no failures. Package passes were WPL 149, agent 870, AI 631, book compiler
|
|
48
|
+
22, coding-agent 5,653, protocol 126 and TUI 730. Vitest workers were bounded at
|
|
49
|
+
four; the TUI package used its Node test runner.
|
|
50
|
+
- Seven-package npm pack dry runs and release-surface checks passed. They checked
|
|
51
|
+
package version/file metadata without publication or lifecycle scripts.
|
|
52
|
+
- Gitleaks found no leaks in the staged changes or the six implementation commits
|
|
53
|
+
since v0.98.4. Reports were redacted; this is not clearance for private histories
|
|
54
|
+
outside that range.
|
|
55
|
+
|
|
56
|
+
## Verification repairs
|
|
57
|
+
|
|
58
|
+
The isolated candidate exposed a two-line module-size overrun that the mixed tree
|
|
59
|
+
hid. Passing initial expansion to the failure-card constructor avoids a redundant
|
|
60
|
+
rebuild and keeps the existing size baseline; both initial states are tested.
|
|
61
|
+
|
|
62
|
+
The full suite exposed an execution wrapper spreading a tool's dynamic timeout
|
|
63
|
+
getter into a fixed value. The wrapper now projects tool metadata and forwards
|
|
64
|
+
that getter lazily. Regression checks cover no eager evaluation, timeout changes,
|
|
65
|
+
stale-context rejection and retained execution metadata.
|
|
66
|
+
|
|
67
|
+
The commit hook had only been tested under plain `sh`, while Husky invokes `sh -e`.
|
|
68
|
+
A changed-file `git diff --quiet` status of 1 therefore stopped valid release
|
|
69
|
+
commits. Capturing that status in an OR-list preserves errexit for real failures;
|
|
70
|
+
all six hook fixtures now run under `sh -e`, including changed paths with spaces,
|
|
71
|
+
index preservation and a fatal git-diff error.
|
|
72
|
+
|
|
73
|
+
An initial test launch resolved to the wrong checkout and was cancelled, not counted
|
|
74
|
+
as candidate evidence. Subsequent launches fix both process cwd and npm prefix.
|
|
75
|
+
An isolated HOME also hid the installed Rust toolchain: explicitly supplying its
|
|
76
|
+
location restored the real cargo diagnostic check without changing the test or
|
|
77
|
+
copying credentials. Neither fixture failure was treated as a product pass.
|
|
78
|
+
|
|
79
|
+
## CI environment recovery
|
|
80
|
+
|
|
81
|
+
The first v0.98.5 tag run built the binaries, but its test step could not find
|
|
82
|
+
`/usr/bin/bwrap`. npm publication was not attempted and GitHub Release creation
|
|
83
|
+
was skipped. The default-branch CI workflows now install `bubblewrap` and run an
|
|
84
|
+
unprivileged namespace probe before the suite. Ubuntu 24.04 subsequently refused
|
|
85
|
+
loopback setup inside the namespace. The runtime-test jobs are pinned to Ubuntu
|
|
86
|
+
22.04 LTS, with the same namespace and capability-drop checks. They do not skip
|
|
87
|
+
verified-run tests, disable host security controls or enable a sandbox fallback.
|
|
88
|
+
The older distribution `fd` lacked `--no-require-git`; CI installs the official
|
|
89
|
+
fd 10.4.2 static archive with SHA-256 verification before extraction and probes
|
|
90
|
+
that option before tests. File-search behavior and tests are not weakened.
|
|
91
|
+
|
|
92
|
+
Recovery dispatches the official workflow from `main` with both `tag` and
|
|
93
|
+
`source_ref` fixed to `v0.98.5`. The release tag and its source commit stay unchanged;
|
|
94
|
+
the existing source/tag equality checks remain mandatory.
|
|
95
|
+
|
|
96
|
+
A source-file fingerprint is not an executed-build attestation. Linux observations
|
|
97
|
+
do not establish behavior on every target platform. CI must validate the exact tag,
|
|
98
|
+
build the six platform archives, run checks/tests, publish all seven npm packages
|
|
99
|
+
and create the GitHub Release. Existing token-based CI authentication is unchanged;
|
|
100
|
+
OIDC/Sigstore provenance is not claimed. Publication is complete only when the main
|
|
101
|
+
tag, GitHub Release and seven npm `latest` values agree.
|
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
# Release audit: v0.99.0
|
|
2
|
+
|
|
3
|
+
기준일: 2026-09-13. 사용자가 배포 준비 결과를 확인한 뒤 즉시 배포를 요청했다.
|
|
4
|
+
이 기록은 후보의 검증 근거이며 GitHub/npm 게시 완료를 미리 선언하지 않는다.
|
|
5
|
+
|
|
6
|
+
## 범위와 변경 로그
|
|
7
|
+
|
|
8
|
+
이전 공개 릴리스 `v0.98.5`는 main의 조상이며, 준비 시점의 GitHub Release와 공개 npm
|
|
9
|
+
`latest` 7개도 0.98.5로 일치했다. 원격 `v0.99.0`이 없음을 확인한 뒤 준비했다.
|
|
10
|
+
|
|
11
|
+
`decee7f157`부터 `ca75f4e5cc`까지의 frontier·Devin·Codex SSE·재시도·TUI·설계 문서와,
|
|
12
|
+
이번에 확정한 `3c7d3b4613`(TB 선택 v2), `b21daf1a76`(TB 감사 v2),
|
|
13
|
+
`003f7b081a`(문서·변경 로그·로컬 캡처 제외)을 포함한다. 마지막 세 단위는 각각
|
|
14
|
+
pre-commit 전체 검사를 통과했고, 선택기 40개·감사기 66개 CLI 검사를 통과했다.
|
|
15
|
+
|
|
16
|
+
TB 출력 계약 변경은 minor 증가로 처리했다. `selectionVersion: 2`의 nullable 예상
|
|
17
|
+
시간·총량과 `omk-tb21-audit-report-2`의 필수 완료시각을 Breaking Changes에 기록했다.
|
|
18
|
+
입력 manifest는 v1을 유지한다. 이전 버전 changelog 본문은 바꾸지 않았으며 다음
|
|
19
|
+
작업용 `[Unreleased]`는 비워 두었다. 새로운 벤치마크 성능이나 SOTA 우위는 주장하지 않는다.
|
|
20
|
+
|
|
21
|
+
## 버전·의존성
|
|
22
|
+
|
|
23
|
+
공개 7개 패키지, root/example manifests, 내부 의존 범위, lockfiles, CLI shrinkwrap,
|
|
24
|
+
book compiler의 `PACKAGE_VERSION`, README 버전 링크와 릴리스 노트를 0.99.0에 맞춘다.
|
|
25
|
+
외부 의존성의 버전·resolved URL·integrity 값은 변경하지 않았다. 모델 카탈로그도
|
|
26
|
+
재생성하지 않았다. 운영자 설정·인증·MCP·스킬 활성 목록은 변경하지 않았다.
|
|
27
|
+
|
|
28
|
+
첫 `version:minor` 실행은 manifest 증가 뒤 아직 이전 버전을 가리키는 내부 의존성을
|
|
29
|
+
npm 출시일 제한에 대조하다 중단됐다. 버전 증가를 재실행하지 않고, 기존
|
|
30
|
+
`sync-versions.js`와 lockfile/설치 동기화 단계만 이어갔다. 제한을 완화하거나 기존
|
|
31
|
+
태그를 이동하지 않았다. 추가적인 registry 패키지 버전 변경은 없음을 대조했다.
|
|
32
|
+
|
|
33
|
+
## 관측한 로컬 검증
|
|
34
|
+
|
|
35
|
+
- 환경: Linux, Node.js 24.19.0, npm 11.14.1.
|
|
36
|
+
- 새 버전의 `npm run build`가 7개 workspace 전체에서 종료 0이었다.
|
|
37
|
+
- `test.sh`를 인증 없는 별도 HOME·최소 환경에서 실행했다. 실제 운영자 auth 파일은
|
|
38
|
+
건드리지 않았다. `taskset`으로 네 CPU에 제한했고 `LIVE_E2E=0`,
|
|
39
|
+
`OMK_NO_LOCAL_LLM=1`, `OMK_OFFLINE=1`을 사용했다.
|
|
40
|
+
- 전체 테스트: **8,279 통과, 852 환경·실계정 조건 skip, 실패 0, 종료 0**.
|
|
41
|
+
WPL 149, agent 870, AI 676, book compiler 22, coding-agent 5,695,
|
|
42
|
+
protocol 137, TUI 730개가 통과했다. 이 수치는 CI 실행 결과가 아니다.
|
|
43
|
+
- 준비 단계의 TB 106개와 전체 guard 378개 통과를 전체 제품 테스트 수에 다시 더하지 않는다.
|
|
44
|
+
- 최종 후보 확인 중 공유 트리에 Devin 요청 코드·테스트·문서와 그 변경 로그가 별도로
|
|
45
|
+
바뀐 것을 발견했다. 이 4개 파일의 후속 변경은 제외하고 `003f7b081a`와 릴리스
|
|
46
|
+
메타데이터만 별도 detached worktree에 옮겼다. 변경 로그의 같은 파일에 섞인 후속
|
|
47
|
+
항목도 원래 커밋의 본문과 대조해 분리했다. 운영자의 변경·인증은 덮어쓰지 않는다.
|
|
48
|
+
- 격리 후보에서도 전체 테스트 8,279개 통과·852개 조건 skip·실패 0을 재확인했다.
|
|
49
|
+
공유 트리 실행과 같은 검사를 중복 합산하지 않았다.
|
|
50
|
+
- 격리 후보의 `npm run check`(Node guard 378개 포함), 7개 pack dry-run·공개 entrypoint
|
|
51
|
+
import, 빌드 CLI의 `--version`·`--help`·`run --help`가 모두 종료 0이었다.
|
|
52
|
+
모든 pack의 버전은 0.99.0이고, 필수 진입점 누락·비공개 상태 경로는 없었다.
|
|
53
|
+
- 후보 diff의 Gitleaks 검사는 완전 redaction과 기존 규칙으로 종료 0, 탐지 0건이었다.
|
|
54
|
+
- `--release`의 stale-worktree guard는 별도 최종 배포 gate다. 후보를 main에 연결하고
|
|
55
|
+
이 작업의 임시 체크아웃을 제거한 뒤 기본 checkout에서 실행한다. guard나 이름을
|
|
56
|
+
바꿔 검사를 피하지 않는다.
|
|
57
|
+
|
|
58
|
+
## 배포 완료 조건
|
|
59
|
+
|
|
60
|
+
기존 `build-binaries.yml`의 태그 기반 경로만 사용한다. 로컬 `npm publish`, 인증 변경,
|
|
61
|
+
새로운 실행 권한·모델 호출은 없다. release source와 tag의 동일 SHA 검사를 유지한다.
|
|
62
|
+
후보는 검토한 명시 경로만 stage하고 staged diff 전체를 확인한 뒤 commit·tag·push한다.
|
|
63
|
+
|
|
64
|
+
공식 workflow가 6개 플랫폼 바이너리 빌드, 검사·테스트, npm 7개 패키지 게시,
|
|
65
|
+
GitHub Release 생성을 완료해야 한다. 최종 판정은 태그의 main 포함,
|
|
66
|
+
GitHub `v0.99.0` Release와 7개 npm `latest`의 일치다. 기존 token 기반 인증을 사용하며
|
|
67
|
+
OIDC/Sigstore provenance는 주장하지 않는다. 실패 시 원인을 확인하고 기존 배포의
|
|
68
|
+
무결성 경계를 유지하며, 완료 전에는 배포 성공이라고 보고하지 않는다.
|
package/docs/run-protocol.md
CHANGED
|
@@ -6,7 +6,7 @@ The OMK Run Protocol defines one versioned contract for task execution and evalu
|
|
|
6
6
|
TaskSpec -> ExecutionAttempt -> Observation -> EvaluationResult -> RuntimeDecision
|
|
7
7
|
```
|
|
8
8
|
|
|
9
|
-
`omk-protocol` owns these records and the pure reducers that connect them. Tool execution, persistence, scheduling, routing, and topology remain outside the package.
|
|
9
|
+
`omk-protocol` owns these records and the pure reducers that connect them. Tool execution, persistence, runtime scheduling, routing, and topology selection remain outside the package.
|
|
10
10
|
|
|
11
11
|
## Implemented scope
|
|
12
12
|
|
|
@@ -24,6 +24,39 @@ The first v1 slice is available under `packages/protocol` with schema version `o
|
|
|
24
24
|
|
|
25
25
|
Every top-level record carries `schemaVersion`. Parsers reject unsupported versions, malformed timestamps, duplicate claim IDs, invalid JSON facts, and empty logical conditions.
|
|
26
26
|
|
|
27
|
+
## Isolated command-run profile
|
|
28
|
+
|
|
29
|
+
`RunContract` and `RunStartCommand` add strict, immutable parsers for the opt-in
|
|
30
|
+
`linux-command-v1` and `linux-scripted-agent-v1` profiles. They pin the input digest,
|
|
31
|
+
write scope, commands, stdout assertions and finite phase/snapshot limits. The
|
|
32
|
+
scripted profile also pins 1–16 approved steps and a 1–32 logical request cap for
|
|
33
|
+
an offline AgentSession reference adapter. Unknown authority/provider fields are
|
|
34
|
+
rejected; the host must separately authorize the exact contract digest.
|
|
35
|
+
`RunResumeCommand` / `parseRunResumeCommand()` pin the contract, candidate and exact
|
|
36
|
+
expected revision/generation. The host may acquire at most three generations under
|
|
37
|
+
`MAX_VERIFIED_RUN_GENERATIONS`; a resume command does not authorize replay of a writer.
|
|
38
|
+
The separate `RunWriterRestartCommand` / `parseRunWriterRestartCommand()` binds
|
|
39
|
+
`baseDigest` to a durable input checkpoint for an explicit local writer restart.
|
|
40
|
+
Recovery commands share command IDs and the generation cap; none resets budgets.
|
|
41
|
+
|
|
42
|
+
`linux-command-dag-v1` adds `RunDagWriter` / `RunDagTask` with 1–16 nodes, explicit
|
|
43
|
+
artifact dependencies, disjoint write scopes and 1–2 preapproved command attempts per
|
|
44
|
+
node. Optional `maxConcurrentTasks` accepts only 1 or 2; omission preserves legacy
|
|
45
|
+
serialization and serial execution. `orderRunDag()` provides deterministic FIFO topological order and
|
|
46
|
+
`runDagAncestors()` includes the full transitive input closure. `RunTaskRetryCommand`
|
|
47
|
+
/ `parseRunTaskRetryCommand()` pins a task selection to the original input and exact
|
|
48
|
+
run revision/generation. These pure contracts never authorize dispatch or authenticate
|
|
49
|
+
cached output; the Coordinator checks stored material and records adoption before reuse.
|
|
50
|
+
|
|
51
|
+
The coding-agent's `RunCoordinator` owns execution and its separate v2 journal.
|
|
52
|
+
It binds native v3 receipt cores to the candidate through a supervisor attestation,
|
|
53
|
+
then supplies authenticated checks to the existing claim-closure reducer.
|
|
54
|
+
This does not replace the task/attempt/evaluation contracts or add execution to
|
|
55
|
+
this package. See [Verified Run](verified-run.md) for the implemented CLI/SDK
|
|
56
|
+
path, candidate recovery and checkpoint-based local writer restart, plus the
|
|
57
|
+
bounded command DAG, eager frontier and selective retry, plus remaining live-model,
|
|
58
|
+
verification-edge, plan amendment and control-surface work.
|
|
59
|
+
|
|
27
60
|
## Durable goal lifecycle
|
|
28
61
|
|
|
29
62
|
A durable goal is working-directory state, not a session-file field or a `TaskSpec`. `/goal <objective>` creates or edits `.omk/goals/current.json`; `/goal` without arguments shows its status and round count.
|
|
@@ -1,5 +1,31 @@
|
|
|
1
1
|
# Runtime Algorithms and Direction
|
|
2
2
|
|
|
3
|
+
## Working-tree shared run budgets
|
|
4
|
+
|
|
5
|
+
The SDK `prompt(..., { runBudget })` path now shares a monotonic deadline and
|
|
6
|
+
logical request/concurrency limits across the active prompt's main stream,
|
|
7
|
+
retries, continuations, and first-party summaries using that stream. Exhaustion
|
|
8
|
+
is a non-retryable `budget_exhausted` termination; snapshots keep outstanding
|
|
9
|
+
streams until terminal metadata arrives. No request/time budget is imposed by
|
|
10
|
+
default. Preflight is now owned even for unbounded prompts; unresolved streams
|
|
11
|
+
block a later prompt instead of being discarded when their budget scope closes.
|
|
12
|
+
See [Shared run budgets](sdk.md#shared-run-budgets-sdk-opt-in) for units, zero
|
|
13
|
+
semantics, cancellation, and uncovered paths. This is not billing enforcement,
|
|
14
|
+
a persisted budget, or a hard process-termination deadline.
|
|
15
|
+
|
|
16
|
+
## Working-tree execution ownership (2026-09-10)
|
|
17
|
+
|
|
18
|
+
The live session now retains registered tool-promise ownership after timeout or
|
|
19
|
+
abort, defers prompt settlement/resource-lease release until actual termination,
|
|
20
|
+
and rejects overlapping prompt admission. Shared permits capture request values,
|
|
21
|
+
honor zero capacity, and wake FIFO followers after head cancellation/expiry.
|
|
22
|
+
The internal lane launcher forwards cancellation and respects zero/heavy caps.
|
|
23
|
+
Independent bash calls now own separate cancellation controllers, and one call's
|
|
24
|
+
completion cannot hide another active call. Core teardown waits use a monotonic
|
|
25
|
+
clock, preserving the existing grace interval across wall-clock changes.
|
|
26
|
+
See [Prompt settlement](sdk.md#prompt-settlement) for tests, compatibility, and
|
|
27
|
+
uncovered paths. These changes do not implement a durable verified-run product.
|
|
28
|
+
|
|
3
29
|
## v0.98.3 release delta (2026-09-06)
|
|
4
30
|
|
|
5
31
|
The explicit advisory SDK requires normal first-party model completion, honors cancellation,
|
|
@@ -128,15 +154,18 @@ Evidence:
|
|
|
128
154
|
- `packages/coding-agent/test/context-budget-selection-policy-version.test.ts`
|
|
129
155
|
- `packages/coding-agent/test/context-budget-cache-disk.test.ts`
|
|
130
156
|
|
|
131
|
-
**Working tree:** non-queued native `xai` requests started through `AgentSession.prompt()` now derive a bounded automatic skill grant from live discovered descriptions after ordinary prompt-template expansion. The selector scores task text separately from camelCase-aware path-to-skill-name signals, excludes explicit-only skills, caps automatic matches at three, and adds `headroom` only under lexical or measured context pressure. `AgentSession.prompt()` merges the result with settings/SDK/bang selections only for that request. Queued steering/follow-up messages reuse the active run's system prompt and do not trigger another selection pass.
|
|
157
|
+
**Working tree:** non-queued native `xai` and `devin` requests started through `AgentSession.prompt()` now derive a bounded automatic skill grant from live discovered descriptions after ordinary prompt-template expansion. The selector scores task text separately from camelCase-aware path-to-skill-name signals, excludes explicit-only skills, caps automatic matches at three, and adds `headroom` only under lexical or measured context pressure. `AgentSession.prompt()` merges the result with settings/SDK/bang selections only for that request. Queued steering/follow-up messages reuse the active run's system prompt and do not trigger another selection pass.
|
|
132
158
|
|
|
133
159
|
Evidence:
|
|
134
160
|
|
|
135
161
|
- `packages/coding-agent/src/core/active-skill-state.ts`
|
|
136
162
|
- `packages/coding-agent/src/core/skill-selector.ts`
|
|
163
|
+
- `packages/coding-agent/src/core/harness-skills.ts`
|
|
137
164
|
- `packages/coding-agent/src/core/grok-harness.ts`
|
|
165
|
+
- `packages/coding-agent/src/core/devin-harness.ts`
|
|
138
166
|
- `packages/coding-agent/src/core/agent-session.ts`
|
|
139
167
|
- `packages/coding-agent/test/grok-active-skills.test.ts`
|
|
168
|
+
- `packages/coding-agent/test/devin-active-skills.test.ts`
|
|
140
169
|
- `packages/coding-agent/test/skill-selector.property.test.ts`
|
|
141
170
|
|
|
142
171
|
**Working tree:** context files now treat their global/local relevance baseline
|