omnius 1.0.695 → 1.0.697
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/api/py-embed.js +4215 -2598
- package/dist/index.js +41242 -35198
- package/dist/library.js +8063 -6080
- package/dist/python-cuda-runtime.js +512 -0
- package/dist/update-worker.js +4250 -2633
- package/docs/DISCOVERY.json +802 -1
- package/docs/DISCOVERY.md +22 -1
- package/docs/guides/long-horizon-feature-workflow.md +42 -0
- package/docs/research/aiwg-long-horizon-feature-workflow.md +94 -0
- package/docs/work-orders/long-horizon-feature-delivery/WO-01-native-research-feature-workflow.md +122 -0
- package/docs/work-orders/runtime-health-remediation/TOOL-QUALITY-2026-09-05.md +84 -0
- package/docs/work-orders/runtime-health-remediation/TRACKER.md +48 -1
- package/docs/work-orders/runtime-health-remediation/WO-30-structured-tool-invocation.md +65 -0
- package/docs/work-orders/runtime-health-remediation/WO-31-complete-runtime-policy.md +62 -0
- package/docs/work-orders/runtime-health-remediation/WO-32-evidence-dependent-steering.md +55 -0
- package/docs/work-orders/runtime-health-remediation/WO-33-file-mutation-transactions.md +55 -0
- package/docs/work-orders/runtime-health-remediation/WO-34-shell-authority-and-results.md +46 -0
- package/docs/work-orders/runtime-health-remediation/WO-35-search-and-exploration-isolation.md +56 -0
- package/docs/work-orders/runtime-health-remediation/WO-36-media-evidence-integrity.md +44 -0
- package/docs/work-orders/runtime-health-remediation/WO-37-browser-and-process-lifecycle.md +86 -0
- package/docs/work-orders/runtime-health-remediation/WO-37-point-localization-cancellation.md +24 -0
- package/docs/work-orders/runtime-health-remediation/WO-38-tool-contract-preservation.md +83 -0
- package/docs/work-orders/runtime-health-remediation/WO-39-telegram-working-indicator.md +76 -0
- package/docs/work-orders/runtime-health-remediation/WO-40-web-content-and-crawl-receipts.md +44 -0
- package/docs/work-orders/runtime-health-remediation/WO-41-completion-evidence-consistency.md +45 -0
- package/docs/work-orders/runtime-health-remediation/WO-42-media-execution-and-configuration.md +79 -0
- package/docs/work-orders/runtime-health-remediation/WO-43-telegram-router-progress-boundary.md +71 -0
- package/docs/work-orders/runtime-health-remediation/WO-44-evidence-backed-terminal-results.md +56 -0
- package/docs/work-orders/runtime-health-remediation/WO-45-native-ollama-tool-contract.md +33 -0
- package/npm-shrinkwrap.json +5 -5
- package/package.json +1 -1
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# WO-30: Tool execution must come from structured calls
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Program:** [September 5 tool-quality and live-behavior follow-up](TOOL-QUALITY-2026-09-05.md)
|
|
6
|
+
|
|
7
|
+
## Observed failure
|
|
8
|
+
|
|
9
|
+
Live descriptive text containing “shell tool interface” triggered two escalating call corrections. Source review additionally found automatic execution of bash/sh/shell/zsh response fences through a direct shellTool.execute path, bypassing normal tool validation, interruption receipts, and mutation accounting.
|
|
10
|
+
|
|
11
|
+
## Evidence and code locations
|
|
12
|
+
|
|
13
|
+
- Installed runtime: Omnius 1.0.695, active run telegram-64ac9937dc7647a8-1788593533269-1.
|
|
14
|
+
- Observation window: 2026-09-05 00:32–00:41 PDT.
|
|
15
|
+
- Live evidence: sibling telegram_test/.omnius/context-window-dumps/ and context/steering-ledger.jsonl; sanitized reproductions must avoid private conversation bodies.
|
|
16
|
+
- packages/orchestrator/src/agenticRunner.ts; packages/orchestrator/tests/edit-transport-guidance.test.ts
|
|
17
|
+
|
|
18
|
+
## Root repair
|
|
19
|
+
|
|
20
|
+
Remove prose/fence-to-execution and tool-name keyword coercion. Keep ordinary assistant prose and examples as text. Execute only actual structured tool calls through the existing admission path. Update the advertised protocol.
|
|
21
|
+
|
|
22
|
+
## Acceptance
|
|
23
|
+
|
|
24
|
+
- [x] A documentation fence executes no shell; descriptive tool names cause no coercion; genuine structured calls still execute and receive ordinary receipts.
|
|
25
|
+
- [x] Reproduce before repair with deterministic tests.
|
|
26
|
+
- [x] Focused regressions, affected typechecks, and workspace build pass.
|
|
27
|
+
- [x] Independent review is resolved and scoped commits delivered to origin/main.
|
|
28
|
+
|
|
29
|
+
## Implementation and verification ledger
|
|
30
|
+
|
|
31
|
+
Removed the entire response-fence execution path, the keyword narration classifier, its escalation counter and obsolete file-write heuristic. Updated the runtime protocol. Added structured-tool-authority.test.ts with four shell-language fences, descriptive tool discussion, and a genuine structured-call control. Replaced the obsolete source assertion in edit-transport-guidance.test.ts.
|
|
32
|
+
|
|
33
|
+
Baseline replay: five of six behavioral tests failed against 835fda0f (all four fences executed; descriptive text caused a coercive correction); structured execution passed. After repair: nine tests across both files passed. Logs: /tmp/omnius-wo30-before.log and /tmp/omnius-wo30-test.log. Aggregate build/review and final delivery are recorded in the program ledger.
|
|
34
|
+
|
|
35
|
+
## Runtime boundary
|
|
36
|
+
|
|
37
|
+
Repair the repository. Preserve live Telegram state, installed package, existing publish/ and docs/DISCOVERY.* changes. No live inference, GPU workloads, service restart, package publication, or external messaging. The user owns publication and later live validation.
|
|
38
|
+
|
|
39
|
+
### Independent review: XML protocol admission
|
|
40
|
+
|
|
41
|
+
Review reproduced a second bypass: parseTextToolCalls promoted XML examples from arbitrary prose into actions, including for native Qwen requests. The parser now requires explicit host opt-in and a complete standalone envelope sequence. Fenced/prose examples remain intact and inert, malformed batches reject as a whole, and each accepted call receives a unique transaction ID. Thirteen model-profile tests pass (/tmp/omnius-wo30-xml.log). Runner provider/stream and explicit JSON text-mode integration is recorded below.
|
|
42
|
+
|
|
43
|
+
### Runner protocol integration
|
|
44
|
+
|
|
45
|
+
Provider adapters now normalize only actual wire `tool_calls`. Canonical runner dispatch selects XML from the host's Hermes profile, or JSON from explicit text mode/the individual tools-unsupported retry. Whole standalone envelopes enter ordinary argument validation, tool accounting, interruption, steering, and completion handling. Native transactions take precedence. Code fences, prose, malformed sequences, and structured-output contracts do not grant text-call authority.
|
|
46
|
+
|
|
47
|
+
Removed the retry's fenced-JSON extraction and direct execution branch. Restored the constructor's omitted `textToolMode` option; explicit text mode advertises the standalone JSON grammar and keeps API tools disabled after exposure refreshes. Streaming uses the same admission rules and preserves inert XML examples. A chunk-ending newline no longer becomes an empty internal-marker prefix, preserving fenced output across arbitrary chunk widths.
|
|
48
|
+
|
|
49
|
+
**110 tests passed across seven suites**, including 38 production/provider authority cases, 34 typed-output tests, 22 evidence-steering tests (primary/brute and crash-point recovery), 13 model-profile tests, and three guidance tests. Orchestrator typecheck passed. Logs: `/tmp/omnius-wo30-final.log` and `/tmp/omnius-wo30-types-final.log`. The final explicit-mode exposure-refresh assertion is checked separately in `/tmp/omnius-wo30-json-final.log`. All tests use synthetic workspaces and mocked transports; no live inference or Telegram requests.
|
|
50
|
+
|
|
51
|
+
Final independent review confirmed whole-payload admission and native-call precedence, and found that text catalogs omitted argument schemas. Both text surfaces now serialize host-prepared definitions with full public parameters and a final profile filter. The updated **40 authority cases passed**, including allowed-schema/denied-tool assertions for both explicit mode and the unsupported-tools retry; final orchestrator typecheck passed. Log: `/tmp/omnius-wo30-schema-final.log`. Review findings are resolved; parent owns aggregate clean build and delivery.
|
|
52
|
+
|
|
53
|
+
Aggregate follow-up: keep hidden think blocks out of the visible provider answer while leaving XML tool examples inert. Removing content-derived invocation had also removed the older incidental reasoning strip; the provider boundary now applies only the dedicated reasoning-strip helper. All 67 tests across provider pool/recovery and the 40-case structured-tool authority matrix passed; /tmp/omnius-wo30-think-recovery.log.
|
|
54
|
+
|
|
55
|
+
## Repository closure — September 5
|
|
56
|
+
|
|
57
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
58
|
+
|
|
59
|
+
Scoped repair commits delivered to origin/main: d2cbec82, e0da0aa6, 4465c9b2, 49f34183. Publication and subsequent live acceptance remain with the user.
|
|
60
|
+
|
|
61
|
+
Source and integration locations:
|
|
62
|
+
|
|
63
|
+
- [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
|
|
64
|
+
- [packages/orchestrator/src/modelProfile.ts](../../../packages/orchestrator/src/modelProfile.ts)
|
|
65
|
+
- [packages/orchestrator/src/typed-model-output.ts](../../../packages/orchestrator/src/typed-model-output.ts)
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# WO-31: Preserve complete typed runtime policy
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Program:** [September 5 tool-quality and live-behavior follow-up](TOOL-QUALITY-2026-09-05.md)
|
|
6
|
+
|
|
7
|
+
## Observed failure
|
|
8
|
+
|
|
9
|
+
The published main prompt is cut at exactly 8,000 characters, mid-sentence. The later medium-model controller and appended agent operating contract never reach the model, despite roughly 80% free context. The separate core reference contract survives.
|
|
10
|
+
|
|
11
|
+
## Evidence and code locations
|
|
12
|
+
|
|
13
|
+
- Installed runtime: Omnius 1.0.695, active run telegram-64ac9937dc7647a8-1788593533269-1.
|
|
14
|
+
- Observation window: 2026-09-05 00:32–00:41 PDT.
|
|
15
|
+
- Live evidence: sibling telegram_test/.omnius/context-window-dumps/ and context/steering-ledger.jsonl; sanitized reproductions must avoid private conversation bodies.
|
|
16
|
+
- packages/orchestrator/src/agenticRunner.ts; packages/orchestrator/src/context-compiler.ts; packages/orchestrator/src/artifactContract.ts
|
|
17
|
+
|
|
18
|
+
## Root repair
|
|
19
|
+
|
|
20
|
+
Preserve runtime-owned policy bodies as complete bounded structural sections through both compiler modes and final projection. Continue bounding untrusted/dynamic system material. Use actual request budgeting; do not silently sever invariant tool instructions.
|
|
21
|
+
|
|
22
|
+
## Acceptance
|
|
23
|
+
|
|
24
|
+
- [x] Large invariant policy reaches final requests intact with headroom; untrusted dynamic blocks remain bounded; tight budgets return an explicit capacity outcome; existing small/medium/large contract tests pass.
|
|
25
|
+
- [x] Reproduce before repair with deterministic tests.
|
|
26
|
+
- [x] Focused regressions, affected typechecks, and workspace build pass.
|
|
27
|
+
- [x] Independent review is resolved and scoped commits delivered to origin/main.
|
|
28
|
+
|
|
29
|
+
## Implementation and verification ledger
|
|
30
|
+
|
|
31
|
+
Runtime-owned policyScope metadata now takes priority over textual controller/receipt examples. Policy bodies split into complete sections (at most 8,000 characters) without losing text. Legacy compaction reserves their full capacity, emits an explicit error if policy alone exceeds its budget, and neither GC nor the signal distiller can trim/drop their content. Active compilation retains final canonical request-budget admission, including protected-overflow refusal. Dynamic project state remains independently bounded.
|
|
32
|
+
|
|
33
|
+
Added runtime-policy-delivery.test.ts: exact section reconstruction, both compiler modes, actual medium-tier backend requests, explicit impossible capacity, and untrusted textual-marker control. Task-replacement production tests also require the full operating contract after retirement. Updated controller selectors to recognize actual envelope lines rather than quoted examples in the newly visible policy.
|
|
34
|
+
|
|
35
|
+
Baseline replay of the actual active backend request test failed against d2cbec82 (operating policy missing); fixed focused set passes 33 distinct tests across five files. Orchestrator typecheck passes. Logs: /tmp/omnius-wo31-before.log, /tmp/omnius-wo31-test.log, /tmp/omnius-wo31-artifact.log, /tmp/omnius-wo31-types.log. Aggregate verification and independent review remain tracked by the program.
|
|
36
|
+
|
|
37
|
+
### Final admission follow-up
|
|
38
|
+
|
|
39
|
+
Independent aggregate testing exposed another downstream loss boundary in `context-admission.ts`: policy metadata was ignored, a trailing system message stopped latest-tool protection, and the emergency projection clipped policy or discarded tool transactions to manufacture an admissible request. Four deterministic admission regressions failed before this repair.
|
|
40
|
+
|
|
41
|
+
Final admission now preserves every typed runtime policy section and matches the latest complete tool batch by transaction IDs across trailing system messages. Emergency admission removes only discardable history and retains protected authority, full source results and their metadata. When that complete payload cannot fit, it returns an explicit rejection with zero admitted output tokens. It no longer clips the current user request or replaces it with invented recovery instructions.
|
|
42
|
+
|
|
43
|
+
The small-source headroom fixture now reserves space for the complete policy: synthetic resident noise was reduced from 1,000 to 300 repetitions, and both actual outbound admission budgets are asserted at or below the 16,000-token window. The explicit 9 KB full read remains intact. Updated admission, policy-delivery and evidence-branch suites pass all 34 tests, including both compiler modes through the final admission method, impossible policy/full-read capacity, and emergency tool-batch preservation. Orchestrator typecheck and scoped diff checks pass. All backends are mocked; no live inference or runtime changes were used.
|
|
44
|
+
|
|
45
|
+
## Runtime boundary
|
|
46
|
+
|
|
47
|
+
Repair the repository. Preserve live Telegram state, installed package, existing publish/ and docs/DISCOVERY.* changes. No live inference, GPU workloads, service restart, package publication, or external messaging. The user owns publication and later live validation.
|
|
48
|
+
|
|
49
|
+
Aggregate recovery fixture reconciliation: the prior synthetic learned ceiling left less capacity than the complete policy itself, so the repaired boundary correctly rejected it. The successful-recovery fixture now supplies a ceiling that can fit the complete policy and asserts byte-for-byte policy preservation, changed request identity, reduced output/prefix and total request fit. Impossible capacities retain separate rejection coverage. All 25 context recovery/admission/runtime-policy tests passed; /tmp/omnius-wo31-recovery.log.
|
|
50
|
+
|
|
51
|
+
## Repository closure — September 5
|
|
52
|
+
|
|
53
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
54
|
+
|
|
55
|
+
Scoped repair commits delivered to origin/main: 8f3134f0, 08577d79, c6ec5040. Publication and subsequent live acceptance remain with the user.
|
|
56
|
+
|
|
57
|
+
Source and integration locations:
|
|
58
|
+
|
|
59
|
+
- [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
|
|
60
|
+
- [packages/orchestrator/src/context-compiler.ts](../../../packages/orchestrator/src/context-compiler.ts)
|
|
61
|
+
- [packages/orchestrator/src/context-admission.ts](../../../packages/orchestrator/src/context-admission.ts)
|
|
62
|
+
- [packages/orchestrator/src/artifactContract.ts](../../../packages/orchestrator/src/artifactContract.ts)
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# WO-32: Resolve task scope after referenced evidence arrives
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Program:** [September 5 tool-quality and live-behavior follow-up](TOOL-QUALITY-2026-09-05.md)
|
|
6
|
+
|
|
7
|
+
## Observed failure
|
|
8
|
+
|
|
9
|
+
The latest explicit switch request was reconciled before its referenced audio was transcribed. Transcription succeeded and subsequent reads switched to boutique-agent-services, but taskEpoch stayed 1, no boundary archive was created, the canonical goal remained the old task projection, and old raw source stayed active. Persisted lifecycle entries omit parsed reconciliation fields, so omitted disposition versus explicit continue cannot be distinguished.
|
|
10
|
+
|
|
11
|
+
## Evidence and code locations
|
|
12
|
+
|
|
13
|
+
- Installed runtime: Omnius 1.0.695, active run telegram-64ac9937dc7647a8-1788593533269-1.
|
|
14
|
+
- Observation window: 2026-09-05 00:32–00:41 PDT.
|
|
15
|
+
- Live evidence: sibling telegram_test/.omnius/context-window-dumps/ and context/steering-ledger.jsonl; sanitized reproductions must avoid private conversation bodies.
|
|
16
|
+
- packages/orchestrator/src/{agenticRunner,steeringIntake,typed-model-output}.ts; task scope production tests
|
|
17
|
+
|
|
18
|
+
## Root repair
|
|
19
|
+
|
|
20
|
+
Make scope decisions explicit and auditable, supporting a bounded evidence-read phase before final reconciliation when the new objective depends on referenced media/source. Preserve the current user request and relevant tool result; prevent unrelated old-plan mutations until scope is resolved. No lexical replacement classifier.
|
|
21
|
+
|
|
22
|
+
## Acceptance
|
|
23
|
+
|
|
24
|
+
- [x] A replacement referring to untranscribed audio can read that evidence and then retire old scope exactly once; additive/status requests remain continuation; reconciliation decisions are persisted and recovery-safe.
|
|
25
|
+
- [x] Reproduce before repair with deterministic tests.
|
|
26
|
+
- [x] Focused regressions, affected typechecks, and workspace build pass.
|
|
27
|
+
- [x] Independent review is resolved and scoped commits delivered to origin/main.
|
|
28
|
+
|
|
29
|
+
## Implementation and verification ledger
|
|
30
|
+
|
|
31
|
+
- Added `resolve_after_evidence` with 1–4 exact registered read-call declarations per phase, at most 3 phases. Names, IDs, raw arguments, parsed read modes, profiles, and single-use tickets constrain admission. Unrelated effects and completion stay gated. The canonical `transcribe_file` exception admits only local-file extraction parameters; general tool metadata remains conservative.
|
|
32
|
+
- Applied reconciliations require explicit disposition. The protected user-authority slot and completion obligation remain unresolved through evidence reads. Final replacement adopts the declared evidence transactions and uses the existing durable scope boundary exactly once; explicit additive continuation keeps the epoch.
|
|
33
|
+
- Parsed decisions are persisted with run/epoch/input correlation for both primary and related inputs. A flushed artifact plus atomic selector stores exact authority, goal, evidence tickets, and transcript even before the first scope boundary.
|
|
34
|
+
- Ticket claims are durable before dispatch. Recovery validates identities and normalized transaction arguments, reports interrupted reads without replay, restores the original goal, and preserves the admission gate. Persistence failures cannot grant in-memory continuation.
|
|
35
|
+
- WO35 integration supplies the full scope to file-exploration notes and rebinds tool state at initial run admission and task epoch changes.
|
|
36
|
+
- **77 tests passed** across steering, typed-output, and scope suites before final durability tightening. **22 tests passed** afterward, including primary/brute evidence→replacement and real mocked `runner.run()` crash-point recovery. Orchestrator `tsc --noEmit` passed.
|
|
37
|
+
- Independent review found two defects in the draft: claim persistence occurred too late, and restored receipt anchors lacked argument matching. Both were repaired and covered by the durability tests. Final parent review confirmed argument matching, epoch-before-checkpoint ordering, fail-closed persistence, and exact ticket admission; no additional concrete defect remained in those paths. No live inference or Telegram requests were used.
|
|
38
|
+
- Parent owns final workspace clean build, aggregate checks, and push.
|
|
39
|
+
|
|
40
|
+
## Runtime boundary
|
|
41
|
+
|
|
42
|
+
Repair the repository. Preserve live Telegram state, installed package, existing publish/ and docs/DISCOVERY.* changes. No live inference, GPU workloads, service restart, package publication, or external messaging. The user owns publication and later live validation.
|
|
43
|
+
|
|
44
|
+
## Repository closure — September 5
|
|
45
|
+
|
|
46
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
47
|
+
|
|
48
|
+
Scoped repair commits delivered to origin/main: 5276b844. Publication and subsequent live acceptance remain with the user.
|
|
49
|
+
|
|
50
|
+
Source and integration locations:
|
|
51
|
+
|
|
52
|
+
- [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
|
|
53
|
+
- [packages/orchestrator/src/steeringIntake.ts](../../../packages/orchestrator/src/steeringIntake.ts)
|
|
54
|
+
- [packages/orchestrator/src/steeringEvidence.ts](../../../packages/orchestrator/src/steeringEvidence.ts)
|
|
55
|
+
- [packages/orchestrator/src/typed-model-output.ts](../../../packages/orchestrator/src/typed-model-output.ts)
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# WO-33: Truthful file mutation transactions
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Scope:** file_write, file_edit, file_patch, batch_edit, notebook_edit, structured_file
|
|
6
|
+
|
|
7
|
+
## Reproduced failures
|
|
8
|
+
|
|
9
|
+
- A second batch destination with denied write permission caused an exception after the first file changed, without a mutation receipt.
|
|
10
|
+
- Two file_edit calls using the same starting hash both succeeded and lost one edit in 8/8 isolated trials.
|
|
11
|
+
- A target and its symlink alias were staged separately; the second write erased the first edit.
|
|
12
|
+
- file_write continued after an existing-file read error and overwrote a writable, unreadable file without overwrite or hash authorization.
|
|
13
|
+
- Invalid supplied hashes were interpreted as absent. Missing old anchors plus replacement text elsewhere falsely produced already-applied success.
|
|
14
|
+
|
|
15
|
+
## Implementation
|
|
16
|
+
|
|
17
|
+
The shared file-mutation boundary resolves symlink targets and serializes cooperating tool instances by canonical path. Every destination is read and validated before committing. Complete UTF-8 bodies and original-content rollback copies are staged beside each destination. Existing files are replaced by rename; new files use exclusive linking of their complete staged body. Hashes and aliases are checked again before each commit.
|
|
18
|
+
|
|
19
|
+
If a later commit fails, previous writes are rolled back when their current bytes still match this transaction. The returned result reports partial mutation and the exact unrestored paths when rollback fails; original-content recovery files are retained and named in the error. Invalid UTF-8, unreadable preimages, invalid hash arguments and unsupported hard links fail closed. All six mutation tools use the same boundary and strict hash guard.
|
|
20
|
+
|
|
21
|
+
An absent old anchor now requires fresh inspection even if replacement text occurs elsewhere. True content-preserving edits still return no-op, and cancelling batch edits no longer emit APPLIED lines.
|
|
22
|
+
|
|
23
|
+
## Guarantees and limits
|
|
24
|
+
|
|
25
|
+
- Serialization covers cooperating calls in this Node process, including different tool instances and symlink aliases. It does not lock independent processes or arbitrary external writers.
|
|
26
|
+
- Individual replacements are atomic. A multi-file batch is staged and rollback-capable, **not crash-atomic**. The process can stop between renames; there is no durable transaction recovery journal.
|
|
27
|
+
- A final pre-commit check reduces external-writer races but cannot provide an operating-system compare-and-swap against unrelated writers.
|
|
28
|
+
- Multiply linked existing files are rejected instead of silently breaking hard-link semantics. Parent directories created for a new file can remain after a failed transaction.
|
|
29
|
+
- Ordinary permission bits are explicitly restored after staging so umask cannot silently narrow an existing mode. Special set-id/sticky bits are rejected. Atomic replacement does not promise preservation of external ACLs, extended attributes or original file ownership.
|
|
30
|
+
- A supplied hash requires an existing preimage; file_write and structured_file cannot recreate a deleted file under its stale hash, including during dry runs.
|
|
31
|
+
|
|
32
|
+
## Verification
|
|
33
|
+
|
|
34
|
+
Hermetic tests exercise concurrent tools, symlink aliases, staging failure, second-rename failure, rollback failure with retained recovery bytes, read-denied overwrite refusal, exclusive creation races, invalid hashes across all six tools, invalid UTF-8 and hard links, and cancelling no-ops.
|
|
35
|
+
|
|
36
|
+
- Final focused run: **112 tests passed** across file-mutation, file-edit/write/patch, batch-edit, notebook-edit and structured-file suites.
|
|
37
|
+
- `pnpm --filter @omnius/execution typecheck`: **passed**.
|
|
38
|
+
- No live inference, service mutation, installed runtime changes or package publication.
|
|
39
|
+
|
|
40
|
+
## Repository closure — September 5
|
|
41
|
+
|
|
42
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
43
|
+
|
|
44
|
+
Scoped repair commits delivered to origin/main: a12d6bd2, be0b126b. Publication and subsequent live acceptance remain with the user.
|
|
45
|
+
|
|
46
|
+
Source and integration locations:
|
|
47
|
+
|
|
48
|
+
- [packages/execution/src/tools/file-mutation.ts](../../../packages/execution/src/tools/file-mutation.ts)
|
|
49
|
+
- [packages/execution/src/tools/edit-metadata.ts](../../../packages/execution/src/tools/edit-metadata.ts)
|
|
50
|
+
- [packages/execution/src/tools/file-write.ts](../../../packages/execution/src/tools/file-write.ts)
|
|
51
|
+
- [packages/execution/src/tools/file-edit.ts](../../../packages/execution/src/tools/file-edit.ts)
|
|
52
|
+
- [packages/execution/src/tools/file-patch.ts](../../../packages/execution/src/tools/file-patch.ts)
|
|
53
|
+
- [packages/execution/src/tools/batch-edit.ts](../../../packages/execution/src/tools/batch-edit.ts)
|
|
54
|
+
- [packages/execution/src/tools/notebook-edit.ts](../../../packages/execution/src/tools/notebook-edit.ts)
|
|
55
|
+
- [packages/execution/src/tools/structured-file.ts](../../../packages/execution/src/tools/structured-file.ts)
|
|
@@ -0,0 +1,46 @@
|
|
|
1
|
+
# WO-34: Shell authority and trustworthy process outcomes
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
|
|
6
|
+
## Reproduced failures
|
|
7
|
+
|
|
8
|
+
`find . -delete`, `git branch -D topic` and a `pwd; python3 ...` writer were classified as read-only/concurrency-safe by command-prefix matching. The runner converts that classification into read-only interruption effect tickets.
|
|
9
|
+
|
|
10
|
+
Printing `exit_code: 0` and then exiting 7 produced a failed tool result with structured `primaryExitCode: 0` and `runnerExitCode: 0`. Those fields were reconstructed from rendered command/stdout text. Displayed SIGPIPE markers could similarly override an unrelated real process failure.
|
|
11
|
+
|
|
12
|
+
The legacy process registry marked recovered PID-only sessions killed without sending any signal. Its remote adapter ignored cwd and attempted to obtain exit status with `wait` from a different shell; its stated wait clamp was not applied.
|
|
13
|
+
|
|
14
|
+
## Repair
|
|
15
|
+
|
|
16
|
+
- Read authority requires one literal command with known read-only semantics. Compound syntax, expansion, interpreter programs, executable rg/find options, unknown options and relevant shell/preprocessor configuration stay effectful. This affects classification, not permission to execute the command.
|
|
17
|
+
- Raw exit code, cwd and timeout state come from spawn callbacks. Exact trailing `echo EXIT=$?` / `EXIT_CODE=$?` forms capture the primary status into the wrapper's private file before emitting display output. Literal single-quoted reporters are ordinary output. Truncation and quoted receipt fields cannot alter primary status.
|
|
18
|
+
- Policy outcomes remain distinct from actual process outcomes. SIGPIPE preview treatment requires an actual observed 141. Tool routing and recovery paths do not invent successful process exits. Elevation uses the actual current cwd.
|
|
19
|
+
- The unused legacy registry is explicitly deprecated toward process-lifecycle/process-async. The environment adapter refuses launch because its API cannot establish owned process and exit receipts. Recovered or remote PID-only sessions cannot be terminated without an owned handle. Owned signals return termination_requested; only observed process exit changes terminal state. Waits now honor their configured clamp.
|
|
20
|
+
|
|
21
|
+
## Scope and limits
|
|
22
|
+
|
|
23
|
+
The classifier deliberately treats git and interpreter commands as effectful because hooks/configuration/code can execute additional commands. It is not a general shell parser and does not claim unknown commands are safe. The private trailing-reporter capture is implemented for the existing POSIX reporter form; unavailable primary observations remain unknown. Elevated exact reporter forms cannot establish the primary subcommand's status and do not receive a fabricated primary success.
|
|
24
|
+
|
|
25
|
+
Repository searches found no production source imports or root public export of tools/process-registry; its local inspection methods remain for compatibility/tests. Its unsupported environment launch path now fails explicitly instead of pretending to supervise remote work.
|
|
26
|
+
|
|
27
|
+
## Verification
|
|
28
|
+
|
|
29
|
+
- Combined shell/classifier/receipt/soft-failure/registry run: **115 tests passed across five suites**.
|
|
30
|
+
- Final classifier expansion: **34 tests passed**, including glob expansion and literal double-quote escaping.
|
|
31
|
+
- Execution package typecheck passed.
|
|
32
|
+
- Regression coverage includes real temporary-file deletion classified effectful, receipt-field collisions, false SIGPIPE, private reporter status, literal reporters, routing without execution, recovered termination refusal, termination request versus observed exit, disabled remote launch and fake-timer wait clamping.
|
|
33
|
+
- Reporter capture beyond the stdout cap passed in the combined run.
|
|
34
|
+
- No live inference, installed runtime changes, service changes or publication.
|
|
35
|
+
|
|
36
|
+
## Repository closure — September 5
|
|
37
|
+
|
|
38
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
39
|
+
|
|
40
|
+
Scoped repair commits delivered to origin/main: a6b6e149. Publication and subsequent live acceptance remain with the user.
|
|
41
|
+
|
|
42
|
+
Source and integration locations:
|
|
43
|
+
|
|
44
|
+
- [packages/execution/src/tools/shell-authority.ts](../../../packages/execution/src/tools/shell-authority.ts)
|
|
45
|
+
- [packages/execution/src/tools/shell.ts](../../../packages/execution/src/tools/shell.ts)
|
|
46
|
+
- [packages/execution/src/tools/process-registry.ts](../../../packages/execution/src/tools/process-registry.ts)
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
# WO-35: Search truth, exploration isolation and source coverage
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
|
|
6
|
+
## Reproduced failures
|
|
7
|
+
|
|
8
|
+
- grep_search and file_explore appended regex text as positional command arguments. Searching for `--version` returned the utility version as successful search evidence.
|
|
9
|
+
- file_explore converted invalid regexes, unavailable paths and failed fallback commands into successful no-match results. It also silently truncated search context.
|
|
10
|
+
- Module-global exploration notes crossed tool instances, working directories and tasks; the runner's unscoped compaction getter could import unrelated findings.
|
|
11
|
+
- file_read admitted fractional ranges and emitted inverted canonical ranges beyond EOF. Requested ends could exceed the actual source while the body was shorter.
|
|
12
|
+
- list_directory omitted entries after 100 without marking incomplete coverage or providing a continuation offset. Failed metadata reads became fabricated zero-byte sizes.
|
|
13
|
+
- find_files advertised path globs, but GNU find's basename-only `-name` treated `**/*.ts` as an impossible basename. During repair, a hermetic fixture also confirmed native Node glob silently treats an unreadable subtree as an empty match set.
|
|
14
|
+
|
|
15
|
+
## Repair
|
|
16
|
+
|
|
17
|
+
1. File exploration delegates text search to GrepSearchTool. Both rg and grep delimit pattern data with `-e` and paths with `--`. Only an unavailable rg executable permits fallback. Invalid regexes, inaccessible paths, stderr overflow and process failures remain failures. Context counts are validated. Configured rg preprocessors are disabled. Explicit large files are no longer silently excluded at 2 MiB; timeout and output caps remain bounded. Partial stdout is marked incomplete, and timeout classification uses process fields rather than pattern text in an error message.
|
|
18
|
+
2. Exploration notes are keyed by canonical working directory and host-bound session, task epoch and owner. Unbound instances have private notes; unscoped export calls return no notes and cannot clear other tasks. Snapshots are defensive copies. A read already in flight keeps its original note collection when the tool is rebound. The latest 128 notes are retained per active scope, with weak registry references allowing abandoned scopes to be collected.
|
|
19
|
+
3. Canonical reads and exploration chunks validate integer ranges before reading; beyond-EOF requests fail without canonical receipts or saved findings. Headers and receipts report the actual selected end. UTF-8 decoding is strict and preserves BOM bytes in source hashes.
|
|
20
|
+
4. Directory inventories sort visible entries and accept validated offset/limit pagination. Capped, selected or grouped inventories carry partial materialization metadata and continuation information. Links are identified as links; missing metadata remains unknown.
|
|
21
|
+
5. File discovery uses strict directory reads with minimatch for actual path/globstar semantics. Unreadable directories fail instead of becoming negative evidence. Directory symlinks and named runtime/dependency directories are skipped. Match, traversal-entry and elapsed-time limits are explicit. A direct minimatch dependency reuses the existing locked 10.2.5 package; installation succeeded offline with a frozen lockfile.
|
|
22
|
+
6. The tool discovery catalog now describes batch edits as staged writes with rollback receipts, removing its stronger atomicity claim.
|
|
23
|
+
|
|
24
|
+
## Runner API
|
|
25
|
+
|
|
26
|
+
`getExploreNotes({ workingDir, sessionId, taskEpoch, ownerId })` and `clearExploreNotes` require the same identity used by `FileExploreTool.bindExecutionScope`. The runner must rebind registered tools on task-epoch transitions. Existing no-argument callers are inert for conservative compatibility. Parent integration handles this consumer and epoch binding; this workorder changes no runner code.
|
|
27
|
+
|
|
28
|
+
## Scope and limits
|
|
29
|
+
|
|
30
|
+
Searches and directory pagination are observations of a live filesystem, not a filesystem snapshot. External changes can move entries between pages. Traversal limits fail with incomplete coverage rather than asserting no files. Search excludes the declared generated/runtime directories, and grep fallback uses extended regex semantics rather than claiming full ripgrep compatibility. File discovery applies full relative-path globs while filename-only patterns match recursively. Working notes are ephemeral, bounded navigation state; normal tool history remains the durable record.
|
|
31
|
+
|
|
32
|
+
Multi-file mutation crash atomicity, cross-process CAS and metadata limits are documented in WO-33. This work does not strengthen those guarantees through discovery text.
|
|
33
|
+
|
|
34
|
+
## Verification
|
|
35
|
+
|
|
36
|
+
- Five focused suites: **115 tests passed** across search, exploration, glob discovery, canonical file reads and directory inventory.
|
|
37
|
+
- Final timeout-text regression: **24 grep tests passed**.
|
|
38
|
+
- Execution package typecheck and scoped diff checks passed.
|
|
39
|
+
- Behavioral fixtures cover literal option-like searches on rg and grep, a disabled executable rg preprocessor, stderr overflow, large-file search, invalid regexes and paths, partial search coverage, concurrent scope rebinding, defensive note snapshots, fractional and beyond-EOF ranges, exact BOM hashes, paginated coverage without duplicates, symlink inventory, globstar/path matching, unreadable subtrees and capped discovery.
|
|
40
|
+
- No inference, network downloads, live runtime changes, service changes or publication.
|
|
41
|
+
|
|
42
|
+
## Repository closure — September 5
|
|
43
|
+
|
|
44
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
45
|
+
|
|
46
|
+
Scoped repair commits delivered to origin/main: 8be70469, 5276b844. Publication and subsequent live acceptance remain with the user.
|
|
47
|
+
|
|
48
|
+
Source and integration locations:
|
|
49
|
+
|
|
50
|
+
- [packages/execution/src/tools/grep-search.ts](../../../packages/execution/src/tools/grep-search.ts)
|
|
51
|
+
- [packages/execution/src/tools/glob-find.ts](../../../packages/execution/src/tools/glob-find.ts)
|
|
52
|
+
- [packages/execution/src/tools/file-explore.ts](../../../packages/execution/src/tools/file-explore.ts)
|
|
53
|
+
- [packages/execution/src/tools/explore-tools.ts](../../../packages/execution/src/tools/explore-tools.ts)
|
|
54
|
+
- [packages/execution/src/tools/file-read.ts](../../../packages/execution/src/tools/file-read.ts)
|
|
55
|
+
- [packages/execution/src/tools/list-directory.ts](../../../packages/execution/src/tools/list-directory.ts)
|
|
56
|
+
- [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# WO-36: Preserve transcription evidence and requested capabilities
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Date:** 2026-09-05
|
|
6
|
+
|
|
7
|
+
## Reproduced failures
|
|
8
|
+
|
|
9
|
+
- Different recordings with the same basename transcribed within one second
|
|
10
|
+
overwrite the same artifacts; an earlier receipt then reads the later text.
|
|
11
|
+
- A backend error object becomes successful `no_speech` evidence. A valid
|
|
12
|
+
segments-only response displays speech while returning empty model content
|
|
13
|
+
and `no_speech` status.
|
|
14
|
+
- The managed Whisper branch ignores a requested `diarize: true` option.
|
|
15
|
+
|
|
16
|
+
## Repair and acceptance
|
|
17
|
+
|
|
18
|
+
- Give each accepted transcription its own artifacts and preserve earlier receipts.
|
|
19
|
+
- Validate backend result shape before persisting or classifying silence; derive
|
|
20
|
+
transcript text from valid segments when needed.
|
|
21
|
+
- Reject unsupported managed diarization before model admission, and do not
|
|
22
|
+
silently use that backend as a diarization fallback.
|
|
23
|
+
- Preserve argv literally when invoking the CLI fallback.
|
|
24
|
+
- Exercise all boundaries with mocked backends and temporary files only.
|
|
25
|
+
|
|
26
|
+
## Verification
|
|
27
|
+
|
|
28
|
+
- `pnpm exec vitest run tests/transcribe-tool-artifacts.test.ts tests/transcribe-python-runtime.test.ts`
|
|
29
|
+
from `packages/execution`: **18 passed**.
|
|
30
|
+
- `pnpm exec tsc --noEmit` from `packages/execution`: passed.
|
|
31
|
+
- The new artifact cases freeze time and verify that each receipt still reads
|
|
32
|
+
its own original text. Failed/malformed results create no evidence artifacts;
|
|
33
|
+
segments-only results preserve speech in model content and storage.
|
|
34
|
+
- No live inference, package installation, service changes, or publication.
|
|
35
|
+
|
|
36
|
+
## Repository closure — September 5
|
|
37
|
+
|
|
38
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
39
|
+
|
|
40
|
+
Scoped repair commits delivered to origin/main: 8d4e5ce5. Publication and subsequent live acceptance remain with the user.
|
|
41
|
+
|
|
42
|
+
Source and integration locations:
|
|
43
|
+
|
|
44
|
+
- [packages/execution/src/tools/transcribe-tool.ts](../../../packages/execution/src/tools/transcribe-tool.ts)
|
|
@@ -0,0 +1,86 @@
|
|
|
1
|
+
# WO-37: Preserve browser ownership and process outcome truth
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Date:** 2026-09-05
|
|
6
|
+
|
|
7
|
+
## Reproduced failures
|
|
8
|
+
|
|
9
|
+
- Separate browser tool instances inherit a module-global session; HTTP 500
|
|
10
|
+
responses with positive JSON become successful actions. Session/action HTTP
|
|
11
|
+
calls lack deadlines, and close can terminate another client's service.
|
|
12
|
+
- A video worker timing out can exit zero and be accepted, or ignore TERM and
|
|
13
|
+
leave the operation pending forever. Output buffers are unbounded.
|
|
14
|
+
- A Carbonyl worker exiting nonzero is accepted when its last stdout line says
|
|
15
|
+
`ok: true`; cancellation does not reliably drain child processes.
|
|
16
|
+
|
|
17
|
+
## Repair and acceptance
|
|
18
|
+
|
|
19
|
+
- Own browser session state per tool instance, close only that session, and
|
|
20
|
+
validate bounded HTTP responses across session and action calls.
|
|
21
|
+
- Connect browser cancellation to outstanding HTTP requests.
|
|
22
|
+
- Reuse the shared process runner for deadlines, bounded output, process-group
|
|
23
|
+
termination, and exit status; preserve progress reporting and diagnostics.
|
|
24
|
+
- Test ownership, transport contradictions, timeouts, cancellation, and worker
|
|
25
|
+
exit contradictions using hermetic fixtures without live models or services.
|
|
26
|
+
|
|
27
|
+
## Verification
|
|
28
|
+
|
|
29
|
+
- Browser session state is instance-owned. Every POST carries its session id,
|
|
30
|
+
including close; closing one session no longer kills the shared service.
|
|
31
|
+
Failed closure retains the handle for truthful recovery.
|
|
32
|
+
- Browser HTTP deadlines cover response bodies, not only headers. JSON shape,
|
|
33
|
+
positive action receipts, and HTTP status are checked independently, and
|
|
34
|
+
active requests honor the tool's cancellation hook.
|
|
35
|
+
- Video and Carbonyl workers use `runProcessBuffer` with bounded output and
|
|
36
|
+
process-group termination. Timeout/abort/overflow errors stay failures even
|
|
37
|
+
if a worker's shutdown handler emits exit zero or positive JSON.
|
|
38
|
+
- `pnpm exec vitest run --maxWorkers=4 --minWorkers=1 tests/browser-action-lifecycle.test.ts tests/browser-action-screenshot.test.ts tests/media-worker-lifecycle.test.ts tests/video-generate.test.ts`
|
|
39
|
+
from `packages/execution`: **36 passed**.
|
|
40
|
+
- `pnpm exec tsc --noEmit` from `packages/execution`: passed.
|
|
41
|
+
- Worker tests use in-memory event emitters and mocked process leases; browser
|
|
42
|
+
tests use mock fetch with live subprocess setup forbidden. No model loading,
|
|
43
|
+
external requests, live browser sessions, or service changes occurred.
|
|
44
|
+
- Parent owns integration and push; publication is excluded.
|
|
45
|
+
|
|
46
|
+
## Ownership review follow-up
|
|
47
|
+
|
|
48
|
+
Review found a remaining launch path that killed every listener on the configured browser port after a capability mismatch. Port occupancy is not proof of process ownership. The launcher now returns a clear incompatible-service failure without invoking shell cleanup or spawning a replacement. A focused regression asserts the existing service is untouched; all 10 browser lifecycle tests passed in /tmp/omnius-wo37-ownership.log.
|
|
49
|
+
|
|
50
|
+
## Nested Playwright ownership and Stop follow-up
|
|
51
|
+
|
|
52
|
+
Carbonyl's text fallback constructs `PlaywrightBrowserTool`. Its prior module-global browser/page/context and diagnostic buffers meant a cancellation hook could close another tool instance's flow. These handles now belong to each tool instance; stateful helpers use that instance's asynchronous scope, and event listeners capture their owning session explicitly. Events from an old closed page cannot enter the replacement page's diagnostics. Same-instance overlapping calls are rejected.
|
|
53
|
+
|
|
54
|
+
`cancel()` closes only the owned page/context/browser, waits for asynchronous cleanup, and prevents a late launch or completed request from becoming successful after Stop. Every layer receives cleanup even if another layer fails. Failed closure retains the affected handle for retry and returns failure. Playwright package/browser installation now uses the shared asynchronous argv process runner with the invocation signal and process-group cleanup. The visual-click lookup receives the same signal and checks cancellation before any DOM fallback or subsequent action; its backend cancellation is implemented in the coordinated vision repair.
|
|
55
|
+
|
|
56
|
+
Review also confirmed two synchronous screenshot conversion shell commands, including interpolation of caller-controlled output paths. Both use one bounded asynchronous Python argv/stdin helper now. Image bytes never enter shell or Python program text; cancellation cannot become a successful image fallback, and empty/non-JPEG conversion output preserves the original PNG instead of being labeled JPEG.
|
|
57
|
+
|
|
58
|
+
Focused verification: **40 tests passed across six suites** (`audio-generation-cancellation`, `audio-generate`, `playwright-browser-cancellation`, `playwright-browser-contract`, `playwright-browser-install`, `playwright-image-conversion`). Execution `tsc --noEmit` and scoped `git diff --check` passed. Browser tests use synthetic objects and conversion stubs; no browser, inference, network, or capture operation ran. Parent owns aggregate review/push. Nested Carbonyl signal wiring is a separate coordinated commit.
|
|
59
|
+
|
|
60
|
+
## Invocation and service-startup cancellation follow-up
|
|
61
|
+
|
|
62
|
+
Independent review found that cancelling startup returned without draining the newly spawned server, asynchronous spawn errors had no handler, repeated readiness probes could extend a nominal 60-second wait to roughly seven minutes, and cleanup addressed a mutable global child without establishing a private process group.
|
|
63
|
+
|
|
64
|
+
- A shared owned-service primitive now retains an immutable child handle, creates a private process group, handles asynchronous spawn failures, and enforces one absolute readiness deadline, including stalled probes. TERM escalates to KILL; an undrained process is reported explicitly. Any service-parent close also terminates the remaining owned group, including unexpected exit and synchronous TERM-close. Parent signal handlers are shared and removed after the final owned service closes.
|
|
65
|
+
- Browser startup has one shared launch promise with separate waiting owners. Cancelling one waiter preserves another waiter's startup; cancelling the final waiter drains only the child created by that launch. Ready shared services and incompatible external listeners are not terminated by an individual tool action.
|
|
66
|
+
- Browser and Carbonyl instances reject overlapping execution. Cancelled health probes retain the known browser session. Carbonyl cancellation reaches setup, Python workers, and its nested Playwright fallback. Browser vision-click passes the same owner signal through point discovery and does not dispatch a click after cancellation.
|
|
67
|
+
- Five browser/service suites passed **30 tests** after the visual-point regression; the broader combined follow-up run passed **87 tests across 13 suites**. Tests use mocked services/processes and synthetic inputs; explicit regressions cover both remaining-group cleanup cases requested in independent review. Execution typecheck and scoped diff checks passed.
|
|
68
|
+
|
|
69
|
+
Carbonyl nested cleanup receipt follow-up: a failed Playwright close now makes the fallback fail and retains the browser on the owning Carbonyl instance. Before admitting another action, that instance retries the retained cleanup; it clears the handle only after a successful close receipt. Cancellation preserves any cleanup error in the returned failure. The focused Carbonyl suite passed **3 tests**, including a failed close, blocked new work during retry failure, and successful ownership release.
|
|
70
|
+
|
|
71
|
+
## Repository closure — September 5
|
|
72
|
+
|
|
73
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
74
|
+
|
|
75
|
+
Scoped repair commits delivered to origin/main: 0d6dcc66, fe3c6487, 52cb2623, 484938c1, 72382afa, 4a072473. Publication and subsequent live acceptance remain with the user.
|
|
76
|
+
|
|
77
|
+
Source and integration locations:
|
|
78
|
+
|
|
79
|
+
- [packages/execution/src/tools/browser-action.ts](../../../packages/execution/src/tools/browser-action.ts)
|
|
80
|
+
- [packages/execution/src/tools/carbonyl-browser.ts](../../../packages/execution/src/tools/carbonyl-browser.ts)
|
|
81
|
+
- [packages/execution/src/tools/playwright-browser.ts](../../../packages/execution/src/tools/playwright-browser.ts)
|
|
82
|
+
- [packages/execution/src/tools/vision.ts](../../../packages/execution/src/tools/vision.ts)
|
|
83
|
+
- [packages/execution/src/owned-service-process.ts](../../../packages/execution/src/owned-service-process.ts)
|
|
84
|
+
- [packages/execution/src/process-async.ts](../../../packages/execution/src/process-async.ts)
|
|
85
|
+
|
|
86
|
+
The [point-localization cancellation follow-up](WO-37-point-localization-cancellation.md) closes the nested vision boundary under this same work order.
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
# WO-37 follow-up: Cancel the active point-localization client
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Parent:** [WO-37 browser and process lifecycle](WO-37-browser-and-process-lifecycle.md)
|
|
5
|
+
|
|
6
|
+
## Confirmed boundary
|
|
7
|
+
|
|
8
|
+
Browser visual clicks await `locateImagePoints`, which previously accepted no caller signal. Its Ollama requests used independent deadlines, its HF point worker and dependency setup blocked in synchronous subprocesses, and the Moondream 0.2 SDK exposes neither a signal nor the underlying HTTP request handle. Closing a browser could leave this work running and permit model fallback after cancellation.
|
|
9
|
+
|
|
10
|
+
## Repair
|
|
11
|
+
|
|
12
|
+
`LocateImagePointsOptions.signal` propagates through Station probing, Ollama generation and response-body waits, installed-model alias discovery, owned pull clients, and owned HF setup/inference processes. The shared process runner bounds output and terminates only each created process group. Abort bypasses candidate fallback and does not mark a model permanently unavailable. HF admission and worker cleanup release their exact lease, including admission that arrives after the caller has stopped. Cancellation during lease release cannot return stale coordinates.
|
|
13
|
+
|
|
14
|
+
The SDK wait is bounded and reports an `AbortError` explaining that its underlying shared operation may continue. The runtime does not kill the shared model/server or alter the cached client to simulate cancellation. Late SDK settlement is observed but cannot publish a result or trigger another backend. Browser owners separately pass their active signal and check cancellation before local fallback or clicking.
|
|
15
|
+
|
|
16
|
+
## Verification
|
|
17
|
+
|
|
18
|
+
Hermetic coverage includes pre-dispatch abort, SDK abort/deadline, stalled HTTP headers/body/alias lookup, HF setup interruption, termination of a synthetic Node worker and a synthetic pull client, temporary-image cleanup, late broker admission, recovery without unavailable-cache poisoning, and cancellation during lease release. Existing vision, HF, desktop-click, and vision-action-loop tests remain covered. Logs: `/tmp/omnius-vision-cancel-final.log` and `/tmp/omnius-vision-cancel-types.log`.
|
|
19
|
+
|
|
20
|
+
Final verification: **79 tests passed across five suites**, including 11 new cancellation cases. Execution `tsc --noEmit` and scoped `git diff --check` passed. Final review checked fallback catches, late admission/release, shared-client behavior, and owned process cleanup; the lease-release race found during review is covered by the final regression.
|
|
21
|
+
|
|
22
|
+
Only synthetic child JavaScript and mocked HTTP/SDK/broker results are used. No Python model code, live inference, GPU workload, Telegram request, service restart, or publication is performed. Orchestrator source remains frozen and outside this follow-up.
|
|
23
|
+
|
|
24
|
+
Repository closure: delivered in 72382afa to origin/main. The final execution aggregate includes these cases; see the [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger). Source locations are [vision.ts](../../../packages/execution/src/tools/vision.ts), [browser-action.ts](../../../packages/execution/src/tools/browser-action.ts) and [playwright-browser.ts](../../../packages/execution/src/tools/playwright-browser.ts). Publication and live acceptance remain with the user.
|
|
@@ -0,0 +1,83 @@
|
|
|
1
|
+
# WO-38: Preserve tool contracts across adapters and policy resolution
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Date:** 2026-09-05 PDT
|
|
6
|
+
|
|
7
|
+
## Observed defects
|
|
8
|
+
|
|
9
|
+
The execution-to-runner adapter discarded `alreadyApplied`,
|
|
10
|
+
`artifactMaterialization`, `documentExtraction`, and `runtimeSession` while
|
|
11
|
+
retaining the success flag. A real already-applied file edit lost its receipt;
|
|
12
|
+
a missing native PDF lost its `not_found` extraction outcome. Partial output
|
|
13
|
+
could consequently pass the runner's canonical-receipt exclusion, and browser
|
|
14
|
+
session changes lost their explicit boundary event.
|
|
15
|
+
|
|
16
|
+
Telegram policy resolution returned shared default allowlist sets. Resolving
|
|
17
|
+
one configuration expanded later policy resolutions without that configuration.
|
|
18
|
+
Blocking an advertised alias also failed to block its canonical tool.
|
|
19
|
+
|
|
20
|
+
The runner separately validated Zod input and discarded the parsed value.
|
|
21
|
+
For `{ count: '3', extra: 'strip' }`, a schema coercing count, defaulting mode,
|
|
22
|
+
and stripping unknown fields still executed the original input.
|
|
23
|
+
|
|
24
|
+
## Scope and implementation
|
|
25
|
+
|
|
26
|
+
- Preserve structured result receipts in both direct and streaming adaptation.
|
|
27
|
+
- Give each policy resolution independent sets; honor blocked advertised aliases.
|
|
28
|
+
- Pass parsed runner input consistently to execution, custom validation,
|
|
29
|
+
fingerprints, concurrency metadata, and path/mutation guards. A transformed
|
|
30
|
+
path must not bypass a guard previously applied to an untransformed path.
|
|
31
|
+
- Keep permissive admin filesystem access and existing public/group restrictions.
|
|
32
|
+
|
|
33
|
+
## Verification
|
|
34
|
+
|
|
35
|
+
Hermetic audit reproduced the first two concrete native-tool losses, policy
|
|
36
|
+
cross-call leakage, and discarded schema defaults/coercion with zero network
|
|
37
|
+
attempts. Fixtures contain synthetic content and live only in temporary dirs.
|
|
38
|
+
|
|
39
|
+
Regression coverage: `packages/cli/tests/tool-contract-preservation.test.ts`
|
|
40
|
+
checks native edit/PDF results, canonical exclusion of partial results through
|
|
41
|
+
both stream and direct adapters, browser-session event generation, policy
|
|
42
|
+
isolation, and advertised alias blocking. Existing adapter-binding and Telegram
|
|
43
|
+
media suites must remain green. Runner evidence is recorded by the parent.
|
|
44
|
+
|
|
45
|
+
- 2026-09-05: **55 tests passed** across tool-contract-preservation (9),
|
|
46
|
+
tool-adapter-binding (2), telegram-admin-media-aliases (33), and
|
|
47
|
+
telegram-public-media (11), with the CLI hermetic network boundary active.
|
|
48
|
+
- CLI `tsc --noEmit` and scoped `git diff --check` passed.
|
|
49
|
+
- Runner preparation now occurs before response deduplication, concurrency
|
|
50
|
+
classification, guards, and dispatch. Each native call retains its parsed
|
|
51
|
+
object so non-idempotent transformations are not repeated at execution.
|
|
52
|
+
Direct API, voice-relay, and model-text calls execute parsed objects too.
|
|
53
|
+
- Runner regression: **206 tests passed** across `tool-schema-execution` (7),
|
|
54
|
+
`agenticRunner` (175), and `agenticRunner-context-behavior` (24). Coverage
|
|
55
|
+
includes custom validation, concurrency metadata, normalized fingerprints,
|
|
56
|
+
rejection of non-object transform output, and a transformed path rejected by
|
|
57
|
+
the existing overwrite guard with the user fixture unchanged.
|
|
58
|
+
- Orchestrator `tsc --noEmit` and scoped `git diff --check` passed.
|
|
59
|
+
|
|
60
|
+
## Delivery boundary
|
|
61
|
+
|
|
62
|
+
Scoped source/test/document commits only. No publication, live inference,
|
|
63
|
+
Telegram requests, service restart, or runtime-state repair in this order.
|
|
64
|
+
|
|
65
|
+
## Streaming wrapper and integration review
|
|
66
|
+
|
|
67
|
+
Native streaming previously bypassed the adapter's execution wrapper, allowing middleware admission or result transformation to be skipped. Streaming now runs inside that same wrapper and relays native chunks through a single pending slot. Backpressure prevents unbounded buffering. Wrapper failure remains authoritative after streamed progress, and wrapper timeout or consumer closure cancels and drains the owned native generator.
|
|
68
|
+
|
|
69
|
+
Independent review reproduced two cleanup defects in the initial relay: asynchronous generator cleanup could outlive consumer close, and wrapper failure could leave the producer running. Both are repaired and covered by tool-adapter-stream-lifecycle.test.ts. All 62 tests across five adapter/Telegram media suites passed in /tmp/omnius-wo38-wrapper.log. Native method binding, live shell chunks, process receipts, partial materialization, admission refusal and post-invocation failure are covered.
|
|
70
|
+
|
|
71
|
+
The native edit fixture now expects a proven no-op only when the exact old anchor is present. Missing anchors remain failures even when replacement text occurs elsewhere. A synthetic idempotent receipt separately verifies alreadyApplied propagation; this preserves WO-33's removal of unsafe inferred edit success.
|
|
72
|
+
|
|
73
|
+
## Repository closure — September 5
|
|
74
|
+
|
|
75
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
76
|
+
|
|
77
|
+
Scoped repair commits delivered to origin/main: e4598f71, cae592bc, 48fd2410. Publication and subsequent live acceptance remain with the user.
|
|
78
|
+
|
|
79
|
+
Source and integration locations:
|
|
80
|
+
|
|
81
|
+
- [packages/cli/src/tui/tool-adapter.ts](../../../packages/cli/src/tui/tool-adapter.ts)
|
|
82
|
+
- [packages/cli/src/tui/tool-policy.ts](../../../packages/cli/src/tui/tool-policy.ts)
|
|
83
|
+
- [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
|