omnius 1.0.695 → 1.0.696
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/api/py-embed.js +4215 -2598
- package/dist/index.js +21372 -18080
- package/dist/library.js +4564 -2912
- package/dist/python-cuda-runtime.js +512 -0
- package/dist/update-worker.js +4250 -2633
- package/docs/DISCOVERY.json +577 -1
- package/docs/DISCOVERY.md +16 -1
- package/docs/work-orders/runtime-health-remediation/TOOL-QUALITY-2026-09-05.md +84 -0
- package/docs/work-orders/runtime-health-remediation/TRACKER.md +23 -1
- package/docs/work-orders/runtime-health-remediation/WO-30-structured-tool-invocation.md +65 -0
- package/docs/work-orders/runtime-health-remediation/WO-31-complete-runtime-policy.md +62 -0
- package/docs/work-orders/runtime-health-remediation/WO-32-evidence-dependent-steering.md +55 -0
- package/docs/work-orders/runtime-health-remediation/WO-33-file-mutation-transactions.md +55 -0
- package/docs/work-orders/runtime-health-remediation/WO-34-shell-authority-and-results.md +46 -0
- package/docs/work-orders/runtime-health-remediation/WO-35-search-and-exploration-isolation.md +56 -0
- package/docs/work-orders/runtime-health-remediation/WO-36-media-evidence-integrity.md +44 -0
- package/docs/work-orders/runtime-health-remediation/WO-37-browser-and-process-lifecycle.md +86 -0
- package/docs/work-orders/runtime-health-remediation/WO-37-point-localization-cancellation.md +24 -0
- package/docs/work-orders/runtime-health-remediation/WO-38-tool-contract-preservation.md +83 -0
- package/docs/work-orders/runtime-health-remediation/WO-39-telegram-working-indicator.md +76 -0
- package/docs/work-orders/runtime-health-remediation/WO-40-web-content-and-crawl-receipts.md +44 -0
- package/docs/work-orders/runtime-health-remediation/WO-41-completion-evidence-consistency.md +45 -0
- package/docs/work-orders/runtime-health-remediation/WO-42-media-execution-and-configuration.md +79 -0
- package/npm-shrinkwrap.json +5 -5
- package/package.json +1 -1
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
# September 5 tool-quality and live-behavior follow-up
|
|
2
|
+
|
|
3
|
+
**Status:** WO-30 through WO-42 complete in repository; full verification passed; delivered to origin/main for user publication
|
|
4
|
+
**Repository baseline:** 501e8394 on origin/main
|
|
5
|
+
**Verified source head:** 7005aa41; subsequent closure changes are documentation only
|
|
6
|
+
**Observed running package:** 1.0.695, verified from its actual executable/package path
|
|
7
|
+
**User scope:** monitor live behavior, audit tool implementations, remedy demonstrated poor practices, keep Telegram typing active, and track durable work orders to completion. Publication belongs to the user.
|
|
8
|
+
|
|
9
|
+
## Live observations
|
|
10
|
+
|
|
11
|
+
The observed run, telegram-64ac9937dc7647a8-1788593533269-1, ran from 00:32:13 through 00:58:09 PDT on September 5. Polling and spool processing remained healthy with zero consecutive polling failures. It completed at task epoch 1 without ambiguous interruption effects.
|
|
12
|
+
|
|
13
|
+
The previous compiler stall did not recur. Raw discovery exceeded 48,000 characters while the request retained substantial context headroom, and no memory compiler request was made. Two early trajectory-grounding calls took about 55 seconds combined. The media alias repair worked: transcribe_file resolved message_id:2755 and returned usable evidence in about 4.4 seconds.
|
|
14
|
+
|
|
15
|
+
User steering redirected work into boutique-agent-services, and subsequent reads and edits followed that target. However, the run retained epoch 1, the old goal/source context, and no scope archive. This exposed the need to defer a scope decision until referenced media has actually been read. The outgoing runtime system prompt was also clipped mid-section despite headroom, and innocent descriptions of tool interfaces triggered corrective feedback.
|
|
16
|
+
|
|
17
|
+
The run changed boutique/src/executor.ts and boutique/src/index.ts, ran checks and probes, and delivered a final report. Its terminal record reported completed/ready while its audit ledger retained verification_missing. Several check commands used trailing status-reporting echoes; a zero wrapper status alone does not establish verifier success. Completion obligations and verification coverage are distinct: explicit required checks gate completion, while observed gaps must remain visible in terminal audit data.
|
|
18
|
+
|
|
19
|
+
## Work orders
|
|
20
|
+
|
|
21
|
+
| Order | Confirmed issue and repair area | Status | Scoped delivery to origin/main |
|
|
22
|
+
| --- | --- | --- | --- |
|
|
23
|
+
| [WO-30](WO-30-structured-tool-invocation.md) | Stop prose/fences/XML examples becoming actions; require a host-selected whole tool protocol | complete | d2cbec82, e0da0aa6, 4465c9b2, 49f34183 |
|
|
24
|
+
| [WO-31](WO-31-complete-runtime-policy.md) | Preserve complete typed runtime policy and current evidence through final admission | complete | 8f3134f0, 08577d79, c6ec5040 |
|
|
25
|
+
| [WO-32](WO-32-evidence-dependent-steering.md) | Durable bounded evidence reads before scope decisions, exact tickets, recovery and epoch retirement | complete | 5276b844 |
|
|
26
|
+
| [WO-33](WO-33-file-mutation-transactions.md) | Guard file replacement, hashes, aliases, concurrency, modes and rollback receipts | complete | a12d6bd2, be0b126b |
|
|
27
|
+
| [WO-34](WO-34-shell-authority-and-results.md) | Conservative shell authority and typed process outcome receipts | complete | a6b6e149 |
|
|
28
|
+
| [WO-35](WO-35-search-and-exploration-isolation.md) | Search options/errors, glob semantics, scoped exploration notes and truthful read/list coverage | complete | 8be70469, 5276b844 |
|
|
29
|
+
| [WO-36](WO-36-media-evidence-integrity.md) | Unique transcript identities, validated backend results and explicit diarization support | complete | 8d4e5ce5 |
|
|
30
|
+
| [WO-37](WO-37-browser-and-process-lifecycle.md) | Session/process ownership, bounded transport and workers, cancellation and startup cleanup | complete | 0d6dcc66, fe3c6487, 52cb2623, 484938c1, 72382afa, 4a072473 |
|
|
31
|
+
| [WO-38](WO-38-tool-contract-preservation.md) | Preserve typed results, parsed inputs, isolated policies and streaming execution wrappers | complete | e4598f71, cae592bc, 48fd2410 |
|
|
32
|
+
| [WO-39](WO-39-telegram-working-indicator.md) | Keep three-second typing active throughout DM/group/topic work and final delivery | complete | 53216239, c17ae247 |
|
|
33
|
+
| [WO-40](WO-40-web-content-and-crawl-receipts.md) | Preserve plain/JSON content, honest cache metadata and validated crawl receipts | complete | 9f96cfcb |
|
|
34
|
+
| [WO-41](WO-41-completion-evidence-consistency.md) | Separate verification audit coverage from explicit completion authority and typed expected effects | complete | 6ee251c3 |
|
|
35
|
+
| [WO-42](WO-42-media-execution-and-configuration.md) | Literal capture arguments, isolated transcription setup, bounded media lifetimes and cancellation | complete | 7bc9106d, 52cb2623, b82f88a9, ed341587, 793c5302 |
|
|
36
|
+
|
|
37
|
+
The [point-localization cancellation follow-up](WO-37-point-localization-cancellation.md) is part of WO-37. All independent review findings are resolved; each order includes concrete source locations, acceptance evidence and limitations.
|
|
38
|
+
|
|
39
|
+
## Review coverage and limits
|
|
40
|
+
|
|
41
|
+
The starting execution tool directory inventory contained 143 source files, including helpers. Deep behavioral review covered filesystem mutation, shell/process management, read/search/exploration, transcription, browser/video/capture, web fetch/crawl, registration/adapters, and runner admission/context/steering. Inventory coverage is distinct from behavioral proof; this is not a claim that every tool and external service is proven correct.
|
|
42
|
+
|
|
43
|
+
Tests use temporary fixtures and mocked inference, Telegram, browser and media transports. This work does not publish, send real Telegram messages, start models, alter services, or repair the monitored workspace in place. The observed live behavior remains evidence for installed 1.0.695; new repository repairs require the user's publication and subsequent field test. A final read-only check at 02:24 PDT still found the same latest completion record, last updated at 00:58:09 PDT; this is not a live acceptance test of the new source.
|
|
44
|
+
|
|
45
|
+
File mutations provide process-local ownership, guarded per-file atomic replacement and explicit rollback diagnostics. They do not claim cross-process compare-and-swap or crash-atomic multi-file transactions. Owned child processes and local browser handles are canceled and drained. Shared Comfy workflows or a Moondream SDK operation can continue after client cancellation where the backend provides no scoped cancellation API; returned diagnostics say so. The implementation does not kill another client's shared service to manufacture a successful Stop receipt.
|
|
46
|
+
|
|
47
|
+
## Delivery gates
|
|
48
|
+
|
|
49
|
+
- [x] Consolidated audit findings have reproductions and work orders.
|
|
50
|
+
- [x] Every confirmed defect in this pass is repaired and independently reviewed.
|
|
51
|
+
- [x] Affected regression suites and clean workspace build pass.
|
|
52
|
+
- [x] Work orders record scope, evidence, results and scoped commits delivered to origin/main.
|
|
53
|
+
- [x] Handoff distinguishes repository verification from unpublished/live behavior.
|
|
54
|
+
- Publication and the live Telegram canary belong to the user after this handoff.
|
|
55
|
+
|
|
56
|
+
## Verification ledger
|
|
57
|
+
|
|
58
|
+
Final verification covers source through 7005aa41. Tests use the repository's hermetic network boundary, mock external systems, and use temporary synthetic workers for process-lifecycle cases.
|
|
59
|
+
|
|
60
|
+
| Verification | Result | Local execution log |
|
|
61
|
+
| --- | --- | --- |
|
|
62
|
+
| Complete orchestrator suite, SQLite tests enabled | 215 suites; 2,624 passed, 1 skipped | /tmp/omnius-tool-quality-orchestrator-verified.log |
|
|
63
|
+
| Complete execution suite after clean build | 160 suites; 1,747 passed, 3 skipped | /tmp/omnius-tool-quality-execution-verified.log |
|
|
64
|
+
| Complete CLI suite after clean build | 266 suites; 2,514 passed | /tmp/omnius-tool-quality-cli-verified.log |
|
|
65
|
+
| Clean all workspace packages, remove residual TypeScript build caches, rebuild | All 11 workspace packages passed | /tmp/omnius-tool-quality-clean-build.log |
|
|
66
|
+
| CUDA preparation helper package entry and artifact policy | Actual source build configuration produced an importable 20,494-byte temporary module, no sourcemap; 16 policy tests passed | /tmp/omnius-cuda-worker-package.log; /tmp/omnius-cuda-worker-package-tests.log |
|
|
67
|
+
|
|
68
|
+
**Total: 641 passing suites, 6,885 passing tests, 4 skipped tests, zero failures.** The full CLI suite includes web-chat-transport-lifecycle, web-ui-client-runtime and web-ui-script. Focused review and before-repair reproductions remain recorded in each work order. Test commands used vitest run --maxWorkers=4 --minWorkers=1 in the respective workspace; the orchestrator run additionally set OMNIUS_SQLITE_TESTS=1. Tools: Node 24.14.0, pnpm 9.15.4, npm 11.9.0.
|
|
69
|
+
|
|
70
|
+
Aggregate review corrected an initial completion overreach: generic mutation alone does not impose mandatory verification; 6ee251c3 preserves audit coverage separately from explicit completion obligations. The full-policy recovery and native-edit fixtures now assert the intended full-policy and exact-anchor contracts. Provider recovery also retains reasoning stripping without interpreting content as tool authority. All affected full suites passed after these corrections.
|
|
71
|
+
|
|
72
|
+
The standalone CUDA preparation helper is now required by both package audits and built by the normal publish script. Verification used temporary output because the user's publish staging directory already had unrelated changes. A complete publication tarball was not rebuilt or published during this pass; the user must run the existing clean-build/bundle/pack publish SOP.
|
|
73
|
+
|
|
74
|
+
## User publication and live acceptance
|
|
75
|
+
|
|
76
|
+
After publishing and running the new package, verify the installed executable's actual version before attributing behavior to these repairs.
|
|
77
|
+
|
|
78
|
+
1. In a DM, start work that includes media preparation or a silent tool interval of at least 18 seconds. Typing should refresh every three seconds through native drafts, silent work and final delivery.
|
|
79
|
+
2. Repeat in public and private groups, including a topic. Steer the active run and confirm its one working heartbeat persists; then Stop and confirm activity retires without a stale timer.
|
|
80
|
+
3. Exercise a normal failure and cancellation during browser/media setup. Owned subprocesses must drain; any shared-server continuation limitation must be explicit.
|
|
81
|
+
4. Switch objectives through referenced audio: evidence is read before the explicit scope decision, old scope retires once, and the selected evidence survives recovery.
|
|
82
|
+
5. Inspect actual outgoing context for complete runtime policy and retained source evidence. A capacity refusal must be explicit, and verification audit gaps must remain separate from configured completion requirements.
|
|
83
|
+
|
|
84
|
+
These are future field checks, not completed live evidence. No installed package, Telegram service or monitored workspace was changed in this repair pass.
|
|
@@ -2,7 +2,29 @@
|
|
|
2
2
|
|
|
3
3
|
**Authority:** canonical granular tracker for RHR-2026-09-02
|
|
4
4
|
**Checked-item rule:** code, focused tests, and named evidence must all exist
|
|
5
|
-
**Last reconciled:** 2026-09-
|
|
5
|
+
**Last reconciled:** 2026-09-05
|
|
6
|
+
|
|
7
|
+
## September 5 tool-quality and live-behavior follow-up
|
|
8
|
+
|
|
9
|
+
Current follow-up authority: [TOOL-QUALITY-2026-09-05.md](TOOL-QUALITY-2026-09-05.md).
|
|
10
|
+
|
|
11
|
+
- [x] WO-30 structured tool invocation and prose safety.
|
|
12
|
+
- [x] WO-31 complete runtime policy delivery.
|
|
13
|
+
- [x] WO-32 evidence-dependent scope reconciliation.
|
|
14
|
+
- [x] Tool-family audit findings reproduced, repaired, and delivered.
|
|
15
|
+
- [x] Aggregate regression/build verification and publication handoff.
|
|
16
|
+
- [x] WO-33: File mutation transactions
|
|
17
|
+
- [x] WO-34: Shell authority and process results
|
|
18
|
+
- [x] WO-35: Search and exploration isolation
|
|
19
|
+
- [x] WO-36: Media evidence integrity
|
|
20
|
+
- [x] WO-37: Browser and process lifecycle
|
|
21
|
+
- [x] WO-38: Tool contract preservation
|
|
22
|
+
- [x] WO-39: Telegram working indicator
|
|
23
|
+
- [x] WO-40: Web content and crawl receipts
|
|
24
|
+
- [x] WO-41: Verification audit and typed completion consistency
|
|
25
|
+
- [x] WO-42: Media execution configuration and cancellation
|
|
26
|
+
|
|
27
|
+
Repository acceptance: all 13 follow-up orders closed; 6,885 tests passed, 4 skipped across 641 suites; clean rebuild passed for all 11 workspace packages. Source head 7005aa41 and scoped repair commits are recorded in the program ledger. Publication and subsequent live Telegram acceptance remain with the user.
|
|
6
28
|
|
|
7
29
|
## September 4 field-review follow-up: WO-25 through WO-29
|
|
8
30
|
|
|
@@ -0,0 +1,65 @@
|
|
|
1
|
+
# WO-30: Tool execution must come from structured calls
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Program:** [September 5 tool-quality and live-behavior follow-up](TOOL-QUALITY-2026-09-05.md)
|
|
6
|
+
|
|
7
|
+
## Observed failure
|
|
8
|
+
|
|
9
|
+
Live descriptive text containing “shell tool interface” triggered two escalating call corrections. Source review additionally found automatic execution of bash/sh/shell/zsh response fences through a direct shellTool.execute path, bypassing normal tool validation, interruption receipts, and mutation accounting.
|
|
10
|
+
|
|
11
|
+
## Evidence and code locations
|
|
12
|
+
|
|
13
|
+
- Installed runtime: Omnius 1.0.695, active run telegram-64ac9937dc7647a8-1788593533269-1.
|
|
14
|
+
- Observation window: 2026-09-05 00:32–00:41 PDT.
|
|
15
|
+
- Live evidence: sibling telegram_test/.omnius/context-window-dumps/ and context/steering-ledger.jsonl; sanitized reproductions must avoid private conversation bodies.
|
|
16
|
+
- packages/orchestrator/src/agenticRunner.ts; packages/orchestrator/tests/edit-transport-guidance.test.ts
|
|
17
|
+
|
|
18
|
+
## Root repair
|
|
19
|
+
|
|
20
|
+
Remove prose/fence-to-execution and tool-name keyword coercion. Keep ordinary assistant prose and examples as text. Execute only actual structured tool calls through the existing admission path. Update the advertised protocol.
|
|
21
|
+
|
|
22
|
+
## Acceptance
|
|
23
|
+
|
|
24
|
+
- [x] A documentation fence executes no shell; descriptive tool names cause no coercion; genuine structured calls still execute and receive ordinary receipts.
|
|
25
|
+
- [x] Reproduce before repair with deterministic tests.
|
|
26
|
+
- [x] Focused regressions, affected typechecks, and workspace build pass.
|
|
27
|
+
- [x] Independent review is resolved and scoped commits delivered to origin/main.
|
|
28
|
+
|
|
29
|
+
## Implementation and verification ledger
|
|
30
|
+
|
|
31
|
+
Removed the entire response-fence execution path, the keyword narration classifier, its escalation counter and obsolete file-write heuristic. Updated the runtime protocol. Added structured-tool-authority.test.ts with four shell-language fences, descriptive tool discussion, and a genuine structured-call control. Replaced the obsolete source assertion in edit-transport-guidance.test.ts.
|
|
32
|
+
|
|
33
|
+
Baseline replay: five of six behavioral tests failed against 835fda0f (all four fences executed; descriptive text caused a coercive correction); structured execution passed. After repair: nine tests across both files passed. Logs: /tmp/omnius-wo30-before.log and /tmp/omnius-wo30-test.log. Aggregate build/review and final delivery are recorded in the program ledger.
|
|
34
|
+
|
|
35
|
+
## Runtime boundary
|
|
36
|
+
|
|
37
|
+
Repair the repository. Preserve live Telegram state, installed package, existing publish/ and docs/DISCOVERY.* changes. No live inference, GPU workloads, service restart, package publication, or external messaging. The user owns publication and later live validation.
|
|
38
|
+
|
|
39
|
+
### Independent review: XML protocol admission
|
|
40
|
+
|
|
41
|
+
Review reproduced a second bypass: parseTextToolCalls promoted XML examples from arbitrary prose into actions, including for native Qwen requests. The parser now requires explicit host opt-in and a complete standalone envelope sequence. Fenced/prose examples remain intact and inert, malformed batches reject as a whole, and each accepted call receives a unique transaction ID. Thirteen model-profile tests pass (/tmp/omnius-wo30-xml.log). Runner provider/stream and explicit JSON text-mode integration is recorded below.
|
|
42
|
+
|
|
43
|
+
### Runner protocol integration
|
|
44
|
+
|
|
45
|
+
Provider adapters now normalize only actual wire `tool_calls`. Canonical runner dispatch selects XML from the host's Hermes profile, or JSON from explicit text mode/the individual tools-unsupported retry. Whole standalone envelopes enter ordinary argument validation, tool accounting, interruption, steering, and completion handling. Native transactions take precedence. Code fences, prose, malformed sequences, and structured-output contracts do not grant text-call authority.
|
|
46
|
+
|
|
47
|
+
Removed the retry's fenced-JSON extraction and direct execution branch. Restored the constructor's omitted `textToolMode` option; explicit text mode advertises the standalone JSON grammar and keeps API tools disabled after exposure refreshes. Streaming uses the same admission rules and preserves inert XML examples. A chunk-ending newline no longer becomes an empty internal-marker prefix, preserving fenced output across arbitrary chunk widths.
|
|
48
|
+
|
|
49
|
+
**110 tests passed across seven suites**, including 38 production/provider authority cases, 34 typed-output tests, 22 evidence-steering tests (primary/brute and crash-point recovery), 13 model-profile tests, and three guidance tests. Orchestrator typecheck passed. Logs: `/tmp/omnius-wo30-final.log` and `/tmp/omnius-wo30-types-final.log`. The final explicit-mode exposure-refresh assertion is checked separately in `/tmp/omnius-wo30-json-final.log`. All tests use synthetic workspaces and mocked transports; no live inference or Telegram requests.
|
|
50
|
+
|
|
51
|
+
Final independent review confirmed whole-payload admission and native-call precedence, and found that text catalogs omitted argument schemas. Both text surfaces now serialize host-prepared definitions with full public parameters and a final profile filter. The updated **40 authority cases passed**, including allowed-schema/denied-tool assertions for both explicit mode and the unsupported-tools retry; final orchestrator typecheck passed. Log: `/tmp/omnius-wo30-schema-final.log`. Review findings are resolved; parent owns aggregate clean build and delivery.
|
|
52
|
+
|
|
53
|
+
Aggregate follow-up: keep hidden think blocks out of the visible provider answer while leaving XML tool examples inert. Removing content-derived invocation had also removed the older incidental reasoning strip; the provider boundary now applies only the dedicated reasoning-strip helper. All 67 tests across provider pool/recovery and the 40-case structured-tool authority matrix passed; /tmp/omnius-wo30-think-recovery.log.
|
|
54
|
+
|
|
55
|
+
## Repository closure — September 5
|
|
56
|
+
|
|
57
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
58
|
+
|
|
59
|
+
Scoped repair commits delivered to origin/main: d2cbec82, e0da0aa6, 4465c9b2, 49f34183. Publication and subsequent live acceptance remain with the user.
|
|
60
|
+
|
|
61
|
+
Source and integration locations:
|
|
62
|
+
|
|
63
|
+
- [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
|
|
64
|
+
- [packages/orchestrator/src/modelProfile.ts](../../../packages/orchestrator/src/modelProfile.ts)
|
|
65
|
+
- [packages/orchestrator/src/typed-model-output.ts](../../../packages/orchestrator/src/typed-model-output.ts)
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# WO-31: Preserve complete typed runtime policy
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Program:** [September 5 tool-quality and live-behavior follow-up](TOOL-QUALITY-2026-09-05.md)
|
|
6
|
+
|
|
7
|
+
## Observed failure
|
|
8
|
+
|
|
9
|
+
The published main prompt is cut at exactly 8,000 characters, mid-sentence. The later medium-model controller and appended agent operating contract never reach the model, despite roughly 80% free context. The separate core reference contract survives.
|
|
10
|
+
|
|
11
|
+
## Evidence and code locations
|
|
12
|
+
|
|
13
|
+
- Installed runtime: Omnius 1.0.695, active run telegram-64ac9937dc7647a8-1788593533269-1.
|
|
14
|
+
- Observation window: 2026-09-05 00:32–00:41 PDT.
|
|
15
|
+
- Live evidence: sibling telegram_test/.omnius/context-window-dumps/ and context/steering-ledger.jsonl; sanitized reproductions must avoid private conversation bodies.
|
|
16
|
+
- packages/orchestrator/src/agenticRunner.ts; packages/orchestrator/src/context-compiler.ts; packages/orchestrator/src/artifactContract.ts
|
|
17
|
+
|
|
18
|
+
## Root repair
|
|
19
|
+
|
|
20
|
+
Preserve runtime-owned policy bodies as complete bounded structural sections through both compiler modes and final projection. Continue bounding untrusted/dynamic system material. Use actual request budgeting; do not silently sever invariant tool instructions.
|
|
21
|
+
|
|
22
|
+
## Acceptance
|
|
23
|
+
|
|
24
|
+
- [x] Large invariant policy reaches final requests intact with headroom; untrusted dynamic blocks remain bounded; tight budgets return an explicit capacity outcome; existing small/medium/large contract tests pass.
|
|
25
|
+
- [x] Reproduce before repair with deterministic tests.
|
|
26
|
+
- [x] Focused regressions, affected typechecks, and workspace build pass.
|
|
27
|
+
- [x] Independent review is resolved and scoped commits delivered to origin/main.
|
|
28
|
+
|
|
29
|
+
## Implementation and verification ledger
|
|
30
|
+
|
|
31
|
+
Runtime-owned policyScope metadata now takes priority over textual controller/receipt examples. Policy bodies split into complete sections (at most 8,000 characters) without losing text. Legacy compaction reserves their full capacity, emits an explicit error if policy alone exceeds its budget, and neither GC nor the signal distiller can trim/drop their content. Active compilation retains final canonical request-budget admission, including protected-overflow refusal. Dynamic project state remains independently bounded.
|
|
32
|
+
|
|
33
|
+
Added runtime-policy-delivery.test.ts: exact section reconstruction, both compiler modes, actual medium-tier backend requests, explicit impossible capacity, and untrusted textual-marker control. Task-replacement production tests also require the full operating contract after retirement. Updated controller selectors to recognize actual envelope lines rather than quoted examples in the newly visible policy.
|
|
34
|
+
|
|
35
|
+
Baseline replay of the actual active backend request test failed against d2cbec82 (operating policy missing); fixed focused set passes 33 distinct tests across five files. Orchestrator typecheck passes. Logs: /tmp/omnius-wo31-before.log, /tmp/omnius-wo31-test.log, /tmp/omnius-wo31-artifact.log, /tmp/omnius-wo31-types.log. Aggregate verification and independent review remain tracked by the program.
|
|
36
|
+
|
|
37
|
+
### Final admission follow-up
|
|
38
|
+
|
|
39
|
+
Independent aggregate testing exposed another downstream loss boundary in `context-admission.ts`: policy metadata was ignored, a trailing system message stopped latest-tool protection, and the emergency projection clipped policy or discarded tool transactions to manufacture an admissible request. Four deterministic admission regressions failed before this repair.
|
|
40
|
+
|
|
41
|
+
Final admission now preserves every typed runtime policy section and matches the latest complete tool batch by transaction IDs across trailing system messages. Emergency admission removes only discardable history and retains protected authority, full source results and their metadata. When that complete payload cannot fit, it returns an explicit rejection with zero admitted output tokens. It no longer clips the current user request or replaces it with invented recovery instructions.
|
|
42
|
+
|
|
43
|
+
The small-source headroom fixture now reserves space for the complete policy: synthetic resident noise was reduced from 1,000 to 300 repetitions, and both actual outbound admission budgets are asserted at or below the 16,000-token window. The explicit 9 KB full read remains intact. Updated admission, policy-delivery and evidence-branch suites pass all 34 tests, including both compiler modes through the final admission method, impossible policy/full-read capacity, and emergency tool-batch preservation. Orchestrator typecheck and scoped diff checks pass. All backends are mocked; no live inference or runtime changes were used.
|
|
44
|
+
|
|
45
|
+
## Runtime boundary
|
|
46
|
+
|
|
47
|
+
Repair the repository. Preserve live Telegram state, installed package, existing publish/ and docs/DISCOVERY.* changes. No live inference, GPU workloads, service restart, package publication, or external messaging. The user owns publication and later live validation.
|
|
48
|
+
|
|
49
|
+
Aggregate recovery fixture reconciliation: the prior synthetic learned ceiling left less capacity than the complete policy itself, so the repaired boundary correctly rejected it. The successful-recovery fixture now supplies a ceiling that can fit the complete policy and asserts byte-for-byte policy preservation, changed request identity, reduced output/prefix and total request fit. Impossible capacities retain separate rejection coverage. All 25 context recovery/admission/runtime-policy tests passed; /tmp/omnius-wo31-recovery.log.
|
|
50
|
+
|
|
51
|
+
## Repository closure — September 5
|
|
52
|
+
|
|
53
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
54
|
+
|
|
55
|
+
Scoped repair commits delivered to origin/main: 8f3134f0, 08577d79, c6ec5040. Publication and subsequent live acceptance remain with the user.
|
|
56
|
+
|
|
57
|
+
Source and integration locations:
|
|
58
|
+
|
|
59
|
+
- [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
|
|
60
|
+
- [packages/orchestrator/src/context-compiler.ts](../../../packages/orchestrator/src/context-compiler.ts)
|
|
61
|
+
- [packages/orchestrator/src/context-admission.ts](../../../packages/orchestrator/src/context-admission.ts)
|
|
62
|
+
- [packages/orchestrator/src/artifactContract.ts](../../../packages/orchestrator/src/artifactContract.ts)
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# WO-32: Resolve task scope after referenced evidence arrives
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Program:** [September 5 tool-quality and live-behavior follow-up](TOOL-QUALITY-2026-09-05.md)
|
|
6
|
+
|
|
7
|
+
## Observed failure
|
|
8
|
+
|
|
9
|
+
The latest explicit switch request was reconciled before its referenced audio was transcribed. Transcription succeeded and subsequent reads switched to boutique-agent-services, but taskEpoch stayed 1, no boundary archive was created, the canonical goal remained the old task projection, and old raw source stayed active. Persisted lifecycle entries omit parsed reconciliation fields, so omitted disposition versus explicit continue cannot be distinguished.
|
|
10
|
+
|
|
11
|
+
## Evidence and code locations
|
|
12
|
+
|
|
13
|
+
- Installed runtime: Omnius 1.0.695, active run telegram-64ac9937dc7647a8-1788593533269-1.
|
|
14
|
+
- Observation window: 2026-09-05 00:32–00:41 PDT.
|
|
15
|
+
- Live evidence: sibling telegram_test/.omnius/context-window-dumps/ and context/steering-ledger.jsonl; sanitized reproductions must avoid private conversation bodies.
|
|
16
|
+
- packages/orchestrator/src/{agenticRunner,steeringIntake,typed-model-output}.ts; task scope production tests
|
|
17
|
+
|
|
18
|
+
## Root repair
|
|
19
|
+
|
|
20
|
+
Make scope decisions explicit and auditable, supporting a bounded evidence-read phase before final reconciliation when the new objective depends on referenced media/source. Preserve the current user request and relevant tool result; prevent unrelated old-plan mutations until scope is resolved. No lexical replacement classifier.
|
|
21
|
+
|
|
22
|
+
## Acceptance
|
|
23
|
+
|
|
24
|
+
- [x] A replacement referring to untranscribed audio can read that evidence and then retire old scope exactly once; additive/status requests remain continuation; reconciliation decisions are persisted and recovery-safe.
|
|
25
|
+
- [x] Reproduce before repair with deterministic tests.
|
|
26
|
+
- [x] Focused regressions, affected typechecks, and workspace build pass.
|
|
27
|
+
- [x] Independent review is resolved and scoped commits delivered to origin/main.
|
|
28
|
+
|
|
29
|
+
## Implementation and verification ledger
|
|
30
|
+
|
|
31
|
+
- Added `resolve_after_evidence` with 1–4 exact registered read-call declarations per phase, at most 3 phases. Names, IDs, raw arguments, parsed read modes, profiles, and single-use tickets constrain admission. Unrelated effects and completion stay gated. The canonical `transcribe_file` exception admits only local-file extraction parameters; general tool metadata remains conservative.
|
|
32
|
+
- Applied reconciliations require explicit disposition. The protected user-authority slot and completion obligation remain unresolved through evidence reads. Final replacement adopts the declared evidence transactions and uses the existing durable scope boundary exactly once; explicit additive continuation keeps the epoch.
|
|
33
|
+
- Parsed decisions are persisted with run/epoch/input correlation for both primary and related inputs. A flushed artifact plus atomic selector stores exact authority, goal, evidence tickets, and transcript even before the first scope boundary.
|
|
34
|
+
- Ticket claims are durable before dispatch. Recovery validates identities and normalized transaction arguments, reports interrupted reads without replay, restores the original goal, and preserves the admission gate. Persistence failures cannot grant in-memory continuation.
|
|
35
|
+
- WO35 integration supplies the full scope to file-exploration notes and rebinds tool state at initial run admission and task epoch changes.
|
|
36
|
+
- **77 tests passed** across steering, typed-output, and scope suites before final durability tightening. **22 tests passed** afterward, including primary/brute evidence→replacement and real mocked `runner.run()` crash-point recovery. Orchestrator `tsc --noEmit` passed.
|
|
37
|
+
- Independent review found two defects in the draft: claim persistence occurred too late, and restored receipt anchors lacked argument matching. Both were repaired and covered by the durability tests. Final parent review confirmed argument matching, epoch-before-checkpoint ordering, fail-closed persistence, and exact ticket admission; no additional concrete defect remained in those paths. No live inference or Telegram requests were used.
|
|
38
|
+
- Parent owns final workspace clean build, aggregate checks, and push.
|
|
39
|
+
|
|
40
|
+
## Runtime boundary
|
|
41
|
+
|
|
42
|
+
Repair the repository. Preserve live Telegram state, installed package, existing publish/ and docs/DISCOVERY.* changes. No live inference, GPU workloads, service restart, package publication, or external messaging. The user owns publication and later live validation.
|
|
43
|
+
|
|
44
|
+
## Repository closure — September 5
|
|
45
|
+
|
|
46
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
47
|
+
|
|
48
|
+
Scoped repair commits delivered to origin/main: 5276b844. Publication and subsequent live acceptance remain with the user.
|
|
49
|
+
|
|
50
|
+
Source and integration locations:
|
|
51
|
+
|
|
52
|
+
- [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
|
|
53
|
+
- [packages/orchestrator/src/steeringIntake.ts](../../../packages/orchestrator/src/steeringIntake.ts)
|
|
54
|
+
- [packages/orchestrator/src/steeringEvidence.ts](../../../packages/orchestrator/src/steeringEvidence.ts)
|
|
55
|
+
- [packages/orchestrator/src/typed-model-output.ts](../../../packages/orchestrator/src/typed-model-output.ts)
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# WO-33: Truthful file mutation transactions
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Scope:** file_write, file_edit, file_patch, batch_edit, notebook_edit, structured_file
|
|
6
|
+
|
|
7
|
+
## Reproduced failures
|
|
8
|
+
|
|
9
|
+
- A second batch destination with denied write permission caused an exception after the first file changed, without a mutation receipt.
|
|
10
|
+
- Two file_edit calls using the same starting hash both succeeded and lost one edit in 8/8 isolated trials.
|
|
11
|
+
- A target and its symlink alias were staged separately; the second write erased the first edit.
|
|
12
|
+
- file_write continued after an existing-file read error and overwrote a writable, unreadable file without overwrite or hash authorization.
|
|
13
|
+
- Invalid supplied hashes were interpreted as absent. Missing old anchors plus replacement text elsewhere falsely produced already-applied success.
|
|
14
|
+
|
|
15
|
+
## Implementation
|
|
16
|
+
|
|
17
|
+
The shared file-mutation boundary resolves symlink targets and serializes cooperating tool instances by canonical path. Every destination is read and validated before committing. Complete UTF-8 bodies and original-content rollback copies are staged beside each destination. Existing files are replaced by rename; new files use exclusive linking of their complete staged body. Hashes and aliases are checked again before each commit.
|
|
18
|
+
|
|
19
|
+
If a later commit fails, previous writes are rolled back when their current bytes still match this transaction. The returned result reports partial mutation and the exact unrestored paths when rollback fails; original-content recovery files are retained and named in the error. Invalid UTF-8, unreadable preimages, invalid hash arguments and unsupported hard links fail closed. All six mutation tools use the same boundary and strict hash guard.
|
|
20
|
+
|
|
21
|
+
An absent old anchor now requires fresh inspection even if replacement text occurs elsewhere. True content-preserving edits still return no-op, and cancelling batch edits no longer emit APPLIED lines.
|
|
22
|
+
|
|
23
|
+
## Guarantees and limits
|
|
24
|
+
|
|
25
|
+
- Serialization covers cooperating calls in this Node process, including different tool instances and symlink aliases. It does not lock independent processes or arbitrary external writers.
|
|
26
|
+
- Individual replacements are atomic. A multi-file batch is staged and rollback-capable, **not crash-atomic**. The process can stop between renames; there is no durable transaction recovery journal.
|
|
27
|
+
- A final pre-commit check reduces external-writer races but cannot provide an operating-system compare-and-swap against unrelated writers.
|
|
28
|
+
- Multiply linked existing files are rejected instead of silently breaking hard-link semantics. Parent directories created for a new file can remain after a failed transaction.
|
|
29
|
+
- Ordinary permission bits are explicitly restored after staging so umask cannot silently narrow an existing mode. Special set-id/sticky bits are rejected. Atomic replacement does not promise preservation of external ACLs, extended attributes or original file ownership.
|
|
30
|
+
- A supplied hash requires an existing preimage; file_write and structured_file cannot recreate a deleted file under its stale hash, including during dry runs.
|
|
31
|
+
|
|
32
|
+
## Verification
|
|
33
|
+
|
|
34
|
+
Hermetic tests exercise concurrent tools, symlink aliases, staging failure, second-rename failure, rollback failure with retained recovery bytes, read-denied overwrite refusal, exclusive creation races, invalid hashes across all six tools, invalid UTF-8 and hard links, and cancelling no-ops.
|
|
35
|
+
|
|
36
|
+
- Final focused run: **112 tests passed** across file-mutation, file-edit/write/patch, batch-edit, notebook-edit and structured-file suites.
|
|
37
|
+
- `pnpm --filter @omnius/execution typecheck`: **passed**.
|
|
38
|
+
- No live inference, service mutation, installed runtime changes or package publication.
|
|
39
|
+
|
|
40
|
+
## Repository closure — September 5
|
|
41
|
+
|
|
42
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
43
|
+
|
|
44
|
+
Scoped repair commits delivered to origin/main: a12d6bd2, be0b126b. Publication and subsequent live acceptance remain with the user.
|
|
45
|
+
|
|
46
|
+
Source and integration locations:
|
|
47
|
+
|
|
48
|
+
- [packages/execution/src/tools/file-mutation.ts](../../../packages/execution/src/tools/file-mutation.ts)
|
|
49
|
+
- [packages/execution/src/tools/edit-metadata.ts](../../../packages/execution/src/tools/edit-metadata.ts)
|
|
50
|
+
- [packages/execution/src/tools/file-write.ts](../../../packages/execution/src/tools/file-write.ts)
|
|
51
|
+
- [packages/execution/src/tools/file-edit.ts](../../../packages/execution/src/tools/file-edit.ts)
|
|
52
|
+
- [packages/execution/src/tools/file-patch.ts](../../../packages/execution/src/tools/file-patch.ts)
|
|
53
|
+
- [packages/execution/src/tools/batch-edit.ts](../../../packages/execution/src/tools/batch-edit.ts)
|
|
54
|
+
- [packages/execution/src/tools/notebook-edit.ts](../../../packages/execution/src/tools/notebook-edit.ts)
|
|
55
|
+
- [packages/execution/src/tools/structured-file.ts](../../../packages/execution/src/tools/structured-file.ts)
|
|
@@ -0,0 +1,46 @@
|
|
|
1
|
+
# WO-34: Shell authority and trustworthy process outcomes
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
|
|
6
|
+
## Reproduced failures
|
|
7
|
+
|
|
8
|
+
`find . -delete`, `git branch -D topic` and a `pwd; python3 ...` writer were classified as read-only/concurrency-safe by command-prefix matching. The runner converts that classification into read-only interruption effect tickets.
|
|
9
|
+
|
|
10
|
+
Printing `exit_code: 0` and then exiting 7 produced a failed tool result with structured `primaryExitCode: 0` and `runnerExitCode: 0`. Those fields were reconstructed from rendered command/stdout text. Displayed SIGPIPE markers could similarly override an unrelated real process failure.
|
|
11
|
+
|
|
12
|
+
The legacy process registry marked recovered PID-only sessions killed without sending any signal. Its remote adapter ignored cwd and attempted to obtain exit status with `wait` from a different shell; its stated wait clamp was not applied.
|
|
13
|
+
|
|
14
|
+
## Repair
|
|
15
|
+
|
|
16
|
+
- Read authority requires one literal command with known read-only semantics. Compound syntax, expansion, interpreter programs, executable rg/find options, unknown options and relevant shell/preprocessor configuration stay effectful. This affects classification, not permission to execute the command.
|
|
17
|
+
- Raw exit code, cwd and timeout state come from spawn callbacks. Exact trailing `echo EXIT=$?` / `EXIT_CODE=$?` forms capture the primary status into the wrapper's private file before emitting display output. Literal single-quoted reporters are ordinary output. Truncation and quoted receipt fields cannot alter primary status.
|
|
18
|
+
- Policy outcomes remain distinct from actual process outcomes. SIGPIPE preview treatment requires an actual observed 141. Tool routing and recovery paths do not invent successful process exits. Elevation uses the actual current cwd.
|
|
19
|
+
- The unused legacy registry is explicitly deprecated toward process-lifecycle/process-async. The environment adapter refuses launch because its API cannot establish owned process and exit receipts. Recovered or remote PID-only sessions cannot be terminated without an owned handle. Owned signals return termination_requested; only observed process exit changes terminal state. Waits now honor their configured clamp.
|
|
20
|
+
|
|
21
|
+
## Scope and limits
|
|
22
|
+
|
|
23
|
+
The classifier deliberately treats git and interpreter commands as effectful because hooks/configuration/code can execute additional commands. It is not a general shell parser and does not claim unknown commands are safe. The private trailing-reporter capture is implemented for the existing POSIX reporter form; unavailable primary observations remain unknown. Elevated exact reporter forms cannot establish the primary subcommand's status and do not receive a fabricated primary success.
|
|
24
|
+
|
|
25
|
+
Repository searches found no production source imports or root public export of tools/process-registry; its local inspection methods remain for compatibility/tests. Its unsupported environment launch path now fails explicitly instead of pretending to supervise remote work.
|
|
26
|
+
|
|
27
|
+
## Verification
|
|
28
|
+
|
|
29
|
+
- Combined shell/classifier/receipt/soft-failure/registry run: **115 tests passed across five suites**.
|
|
30
|
+
- Final classifier expansion: **34 tests passed**, including glob expansion and literal double-quote escaping.
|
|
31
|
+
- Execution package typecheck passed.
|
|
32
|
+
- Regression coverage includes real temporary-file deletion classified effectful, receipt-field collisions, false SIGPIPE, private reporter status, literal reporters, routing without execution, recovered termination refusal, termination request versus observed exit, disabled remote launch and fake-timer wait clamping.
|
|
33
|
+
- Reporter capture beyond the stdout cap passed in the combined run.
|
|
34
|
+
- No live inference, installed runtime changes, service changes or publication.
|
|
35
|
+
|
|
36
|
+
## Repository closure — September 5
|
|
37
|
+
|
|
38
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
39
|
+
|
|
40
|
+
Scoped repair commits delivered to origin/main: a6b6e149. Publication and subsequent live acceptance remain with the user.
|
|
41
|
+
|
|
42
|
+
Source and integration locations:
|
|
43
|
+
|
|
44
|
+
- [packages/execution/src/tools/shell-authority.ts](../../../packages/execution/src/tools/shell-authority.ts)
|
|
45
|
+
- [packages/execution/src/tools/shell.ts](../../../packages/execution/src/tools/shell.ts)
|
|
46
|
+
- [packages/execution/src/tools/process-registry.ts](../../../packages/execution/src/tools/process-registry.ts)
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
# WO-35: Search truth, exploration isolation and source coverage
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
|
|
6
|
+
## Reproduced failures
|
|
7
|
+
|
|
8
|
+
- grep_search and file_explore appended regex text as positional command arguments. Searching for `--version` returned the utility version as successful search evidence.
|
|
9
|
+
- file_explore converted invalid regexes, unavailable paths and failed fallback commands into successful no-match results. It also silently truncated search context.
|
|
10
|
+
- Module-global exploration notes crossed tool instances, working directories and tasks; the runner's unscoped compaction getter could import unrelated findings.
|
|
11
|
+
- file_read admitted fractional ranges and emitted inverted canonical ranges beyond EOF. Requested ends could exceed the actual source while the body was shorter.
|
|
12
|
+
- list_directory omitted entries after 100 without marking incomplete coverage or providing a continuation offset. Failed metadata reads became fabricated zero-byte sizes.
|
|
13
|
+
- find_files advertised path globs, but GNU find's basename-only `-name` treated `**/*.ts` as an impossible basename. During repair, a hermetic fixture also confirmed native Node glob silently treats an unreadable subtree as an empty match set.
|
|
14
|
+
|
|
15
|
+
## Repair
|
|
16
|
+
|
|
17
|
+
1. File exploration delegates text search to GrepSearchTool. Both rg and grep delimit pattern data with `-e` and paths with `--`. Only an unavailable rg executable permits fallback. Invalid regexes, inaccessible paths, stderr overflow and process failures remain failures. Context counts are validated. Configured rg preprocessors are disabled. Explicit large files are no longer silently excluded at 2 MiB; timeout and output caps remain bounded. Partial stdout is marked incomplete, and timeout classification uses process fields rather than pattern text in an error message.
|
|
18
|
+
2. Exploration notes are keyed by canonical working directory and host-bound session, task epoch and owner. Unbound instances have private notes; unscoped export calls return no notes and cannot clear other tasks. Snapshots are defensive copies. A read already in flight keeps its original note collection when the tool is rebound. The latest 128 notes are retained per active scope, with weak registry references allowing abandoned scopes to be collected.
|
|
19
|
+
3. Canonical reads and exploration chunks validate integer ranges before reading; beyond-EOF requests fail without canonical receipts or saved findings. Headers and receipts report the actual selected end. UTF-8 decoding is strict and preserves BOM bytes in source hashes.
|
|
20
|
+
4. Directory inventories sort visible entries and accept validated offset/limit pagination. Capped, selected or grouped inventories carry partial materialization metadata and continuation information. Links are identified as links; missing metadata remains unknown.
|
|
21
|
+
5. File discovery uses strict directory reads with minimatch for actual path/globstar semantics. Unreadable directories fail instead of becoming negative evidence. Directory symlinks and named runtime/dependency directories are skipped. Match, traversal-entry and elapsed-time limits are explicit. A direct minimatch dependency reuses the existing locked 10.2.5 package; installation succeeded offline with a frozen lockfile.
|
|
22
|
+
6. The tool discovery catalog now describes batch edits as staged writes with rollback receipts, removing its stronger atomicity claim.
|
|
23
|
+
|
|
24
|
+
## Runner API
|
|
25
|
+
|
|
26
|
+
`getExploreNotes({ workingDir, sessionId, taskEpoch, ownerId })` and `clearExploreNotes` require the same identity used by `FileExploreTool.bindExecutionScope`. The runner must rebind registered tools on task-epoch transitions. Existing no-argument callers are inert for conservative compatibility. Parent integration handles this consumer and epoch binding; this workorder changes no runner code.
|
|
27
|
+
|
|
28
|
+
## Scope and limits
|
|
29
|
+
|
|
30
|
+
Searches and directory pagination are observations of a live filesystem, not a filesystem snapshot. External changes can move entries between pages. Traversal limits fail with incomplete coverage rather than asserting no files. Search excludes the declared generated/runtime directories, and grep fallback uses extended regex semantics rather than claiming full ripgrep compatibility. File discovery applies full relative-path globs while filename-only patterns match recursively. Working notes are ephemeral, bounded navigation state; normal tool history remains the durable record.
|
|
31
|
+
|
|
32
|
+
Multi-file mutation crash atomicity, cross-process CAS and metadata limits are documented in WO-33. This work does not strengthen those guarantees through discovery text.
|
|
33
|
+
|
|
34
|
+
## Verification
|
|
35
|
+
|
|
36
|
+
- Five focused suites: **115 tests passed** across search, exploration, glob discovery, canonical file reads and directory inventory.
|
|
37
|
+
- Final timeout-text regression: **24 grep tests passed**.
|
|
38
|
+
- Execution package typecheck and scoped diff checks passed.
|
|
39
|
+
- Behavioral fixtures cover literal option-like searches on rg and grep, a disabled executable rg preprocessor, stderr overflow, large-file search, invalid regexes and paths, partial search coverage, concurrent scope rebinding, defensive note snapshots, fractional and beyond-EOF ranges, exact BOM hashes, paginated coverage without duplicates, symlink inventory, globstar/path matching, unreadable subtrees and capped discovery.
|
|
40
|
+
- No inference, network downloads, live runtime changes, service changes or publication.
|
|
41
|
+
|
|
42
|
+
## Repository closure — September 5
|
|
43
|
+
|
|
44
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
45
|
+
|
|
46
|
+
Scoped repair commits delivered to origin/main: 8be70469, 5276b844. Publication and subsequent live acceptance remain with the user.
|
|
47
|
+
|
|
48
|
+
Source and integration locations:
|
|
49
|
+
|
|
50
|
+
- [packages/execution/src/tools/grep-search.ts](../../../packages/execution/src/tools/grep-search.ts)
|
|
51
|
+
- [packages/execution/src/tools/glob-find.ts](../../../packages/execution/src/tools/glob-find.ts)
|
|
52
|
+
- [packages/execution/src/tools/file-explore.ts](../../../packages/execution/src/tools/file-explore.ts)
|
|
53
|
+
- [packages/execution/src/tools/explore-tools.ts](../../../packages/execution/src/tools/explore-tools.ts)
|
|
54
|
+
- [packages/execution/src/tools/file-read.ts](../../../packages/execution/src/tools/file-read.ts)
|
|
55
|
+
- [packages/execution/src/tools/list-directory.ts](../../../packages/execution/src/tools/list-directory.ts)
|
|
56
|
+
- [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# WO-36: Preserve transcription evidence and requested capabilities
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Date:** 2026-09-05
|
|
6
|
+
|
|
7
|
+
## Reproduced failures
|
|
8
|
+
|
|
9
|
+
- Different recordings with the same basename transcribed within one second
|
|
10
|
+
overwrite the same artifacts; an earlier receipt then reads the later text.
|
|
11
|
+
- A backend error object becomes successful `no_speech` evidence. A valid
|
|
12
|
+
segments-only response displays speech while returning empty model content
|
|
13
|
+
and `no_speech` status.
|
|
14
|
+
- The managed Whisper branch ignores a requested `diarize: true` option.
|
|
15
|
+
|
|
16
|
+
## Repair and acceptance
|
|
17
|
+
|
|
18
|
+
- Give each accepted transcription its own artifacts and preserve earlier receipts.
|
|
19
|
+
- Validate backend result shape before persisting or classifying silence; derive
|
|
20
|
+
transcript text from valid segments when needed.
|
|
21
|
+
- Reject unsupported managed diarization before model admission, and do not
|
|
22
|
+
silently use that backend as a diarization fallback.
|
|
23
|
+
- Preserve argv literally when invoking the CLI fallback.
|
|
24
|
+
- Exercise all boundaries with mocked backends and temporary files only.
|
|
25
|
+
|
|
26
|
+
## Verification
|
|
27
|
+
|
|
28
|
+
- `pnpm exec vitest run tests/transcribe-tool-artifacts.test.ts tests/transcribe-python-runtime.test.ts`
|
|
29
|
+
from `packages/execution`: **18 passed**.
|
|
30
|
+
- `pnpm exec tsc --noEmit` from `packages/execution`: passed.
|
|
31
|
+
- The new artifact cases freeze time and verify that each receipt still reads
|
|
32
|
+
its own original text. Failed/malformed results create no evidence artifacts;
|
|
33
|
+
segments-only results preserve speech in model content and storage.
|
|
34
|
+
- No live inference, package installation, service changes, or publication.
|
|
35
|
+
|
|
36
|
+
## Repository closure — September 5
|
|
37
|
+
|
|
38
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
39
|
+
|
|
40
|
+
Scoped repair commits delivered to origin/main: 8d4e5ce5. Publication and subsequent live acceptance remain with the user.
|
|
41
|
+
|
|
42
|
+
Source and integration locations:
|
|
43
|
+
|
|
44
|
+
- [packages/execution/src/tools/transcribe-tool.ts](../../../packages/execution/src/tools/transcribe-tool.ts)
|