omnius 1.0.694 → 1.0.696

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (31) hide show
  1. package/dist/api/py-embed.js +4215 -2598
  2. package/dist/index.js +23327 -18826
  3. package/dist/library.js +4564 -2912
  4. package/dist/python-cuda-runtime.js +512 -0
  5. package/dist/update-worker.js +4250 -2633
  6. package/docs/DISCOVERY.json +809 -1
  7. package/docs/DISCOVERY.md +22 -1
  8. package/docs/work-orders/runtime-health-remediation/TELEGRAM-LONGHAUL-2026-09-04.md +92 -0
  9. package/docs/work-orders/runtime-health-remediation/TOOL-QUALITY-2026-09-05.md +84 -0
  10. package/docs/work-orders/runtime-health-remediation/TRACKER.md +38 -1
  11. package/docs/work-orders/runtime-health-remediation/WO-25-memory-compiler-eligibility-and-audit.md +81 -0
  12. package/docs/work-orders/runtime-health-remediation/WO-26-admin-telegram-media-aliases.md +93 -0
  13. package/docs/work-orders/runtime-health-remediation/WO-27-task-verification-authority.md +140 -0
  14. package/docs/work-orders/runtime-health-remediation/WO-28-task-scope-evidence-retirement.md +82 -0
  15. package/docs/work-orders/runtime-health-remediation/WO-29-retired-failure-projections.md +85 -0
  16. package/docs/work-orders/runtime-health-remediation/WO-30-structured-tool-invocation.md +65 -0
  17. package/docs/work-orders/runtime-health-remediation/WO-31-complete-runtime-policy.md +62 -0
  18. package/docs/work-orders/runtime-health-remediation/WO-32-evidence-dependent-steering.md +55 -0
  19. package/docs/work-orders/runtime-health-remediation/WO-33-file-mutation-transactions.md +55 -0
  20. package/docs/work-orders/runtime-health-remediation/WO-34-shell-authority-and-results.md +46 -0
  21. package/docs/work-orders/runtime-health-remediation/WO-35-search-and-exploration-isolation.md +56 -0
  22. package/docs/work-orders/runtime-health-remediation/WO-36-media-evidence-integrity.md +44 -0
  23. package/docs/work-orders/runtime-health-remediation/WO-37-browser-and-process-lifecycle.md +86 -0
  24. package/docs/work-orders/runtime-health-remediation/WO-37-point-localization-cancellation.md +24 -0
  25. package/docs/work-orders/runtime-health-remediation/WO-38-tool-contract-preservation.md +83 -0
  26. package/docs/work-orders/runtime-health-remediation/WO-39-telegram-working-indicator.md +76 -0
  27. package/docs/work-orders/runtime-health-remediation/WO-40-web-content-and-crawl-receipts.md +44 -0
  28. package/docs/work-orders/runtime-health-remediation/WO-41-completion-evidence-consistency.md +45 -0
  29. package/docs/work-orders/runtime-health-remediation/WO-42-media-execution-and-configuration.md +79 -0
  30. package/npm-shrinkwrap.json +5 -5
  31. package/package.json +1 -1
@@ -0,0 +1,82 @@
1
+ # WO-28: Retire obsolete task evidence after explicit scope changes
2
+
3
+ **Status:** complete; implemented, verified, and pushed to origin/main
4
+ **Priority:** P2
5
+ **Incident date:** 2026-09-04 PDT / 2026-09-05 UTC
6
+ **Program:** [September 4 Telegram long-haul repairs](TELEGRAM-LONGHAUL-2026-09-04.md)
7
+
8
+ ## Observed failure
9
+
10
+ At 20:41 PDT, current turn 21 retained 57,931 characters of previous pentest repository tool bodies (38.6% of all message content), versus 35,088 from the new boutique-services project. The model eventually followed the new focus, but old task material remained in the active frame/history and sustained raw-discovery pressure.
11
+
12
+ ## Evidence
13
+
14
+ Paths are relative to /home/roko/Documents/Projects/Adjacent. Live files are
15
+ observational references; sanitized deterministic fixtures must reproduce the
16
+ failure without copying credentials or private conversation content.
17
+
18
+ - `telegram_test/.omnius/context-window-dumps/2026-09-05T03-41-09-548Z-main-57e9eeba54.json`
19
+
20
+ ## Root cause
21
+
22
+ Telegram deliberately admits normal messages as `context_only`. The typed model reconciliation had no replacement field, and an immediately following voice message overwrote the one unresolved steering slot. Even explicit transport replacements reset only part of the task state: raw messages, evidence-ledger bodies, tool-event context, and task trajectory survived. Recovery also retained an old lifecycle epoch.
23
+
24
+ ## Code locations
25
+
26
+ - `packages/orchestrator/src/agenticRunner.ts`: grouped steering, typed replacement, exact archives, safe checkpoints, task-state reset, request/recovery integration.
27
+ - `packages/orchestrator/src/typed-model-output.ts` and `steeringIntake.ts`: bounded typed `taskDisposition`, `replacementGoal`, retained tool IDs.
28
+ - `packages/orchestrator/src/taskScopeBoundary.ts`: validated run/session/epoch selector and complete tool transactions.
29
+ - `packages/orchestrator/src/interruption-lifecycle.ts`: drained, owner-checked epoch advancement with persistence rollback.
30
+ - `packages/orchestrator/src/contentAddressedArtifactStore.ts`: explicit durable flush for transition archives.
31
+ - `packages/orchestrator/tests/agenticRunner-task-scope.test.ts`, `taskScopeBoundary.test.ts`, `typed-model-output.test.ts`, lifecycle/artifact tests and production request replay.
32
+
33
+ ## Repair design
34
+
35
+ - Trace and repair the typed task-replacement boundary; do not infer replacement from filenames or keyword similarity alone.
36
+ - Retire prior-task raw tool transactions from the next model-visible request after an accepted explicit replacement, preserving their durable audit/receipt references.
37
+ - Retain current user authority, replacement instructions, media reference evidence needed to interpret the new task, and explicitly adopted cross-task dependencies.
38
+ - Do not treat a status question, additive instruction, or priority change as destructive task replacement.
39
+ - Keep canonical goal and evidence selection consistent across continuation/restoration and integrate with WO-25.
40
+
41
+ ## Acceptance checklist
42
+
43
+ - [x] A two-project replacement replay excludes obsolete raw source while preserving current task instructions and required new evidence.
44
+ - [x] Status-only and additive steering preserve relevant active-task context.
45
+ - [x] Retired evidence remains durably retrievable; tool call/result pairs remain valid.
46
+ - [x] Restored continuation uses the same scope boundary without resurrecting obsolete raw history.
47
+ - [x] Integrated compiler/projection replay proves scope retirement does not trigger useless compaction.
48
+ - [x] Focused steering/context tests and orchestrator typecheck pass.
49
+
50
+ ## Implementation and verification log
51
+
52
+ - Initial regression reproduced overwritten text authority, retained obsolete raw history, and unvalidated retained tool IDs (three failing cases).
53
+ - Typed reconciliation batches unresolved text/media and advances scope only for an accepted replacement, with no lexical task classifier. Transport replacement acknowledgement cannot advance twice.
54
+ - Exact immutable transition archives include messages, todos, workboard, completion ledger, and task state. File/directory sync precedes lifecycle epoch advancement and retirement.
55
+ - Later safe transcripts are checkpointed separately so restart retains work performed after replacement. Early recovery selects current goal before startup analysis.
56
+ - Reset covers secondary evidence/context caches, world facts, task trajectory, phase tree, and verifier authority. Explicitly adopted complete tool transactions remain model-visible.
57
+ - Independent review exposed and drove repairs for cache reinjection, structured-state archive loss, restore progress loss, and identical text in later task scope. Final verification and commit evidence follow below.
58
+
59
+ ## Verification result
60
+
61
+ - Runner scope regressions: **13 passed**; pure boundary regressions: **11 passed**.
62
+ - Durable interruption lifecycle: **25 passed**; content-addressed archive store: **10 passed**.
63
+ - Full mocked `runner.run()` replacement replay: **2 passed**, exercising primary and brute-force request paths, grouped text/audio, exact epoch advancement, evidence retirement, and retained core policy.
64
+ - Existing core-policy regressions: **3 passed**, covering small, medium, and large tiers.
65
+ - Compiler/canonical/steering integration regression set passed. Clean workspace build passed for all 11 packages/apps.
66
+ - Final scope reset includes observation and source ledgers, contract registry, local tool cache, world facts, recovery-controller delta, phase tree, and cached prior-task retrieval. Current workspace inventory remains observable.
67
+ - Identical later user replies survive restore: selectors recognize the original archived prefix and stop retirement at the new authority/epoch boundary.
68
+ - Runtime policy has typed metadata and stays in separate bounded system messages. This avoids losing core rules to first-system truncation while retiring old task recall.
69
+ - Full aggregate suite and commit evidence are recorded in the program ledger.
70
+
71
+ ## Delivery
72
+
73
+ - Commit: `24321b30` on `origin/main`.
74
+ - Final archive-directory sync refinement passed both scope suites (**15 tests**) and an additional workspace build. Full aggregate evidence is in the linked program ledger.
75
+
76
+ ## Publication boundary
77
+
78
+ Repository repair and deterministic verification are this order's completion
79
+ boundary. The user will publish the package and then perform live validation.
80
+ No package publication, installed-package replacement, runtime-state repair,
81
+ or service restart is authorized by this order.
82
+
@@ -0,0 +1,85 @@
1
+ # WO-29: Exclude retired failures from active diagnostics
2
+
3
+ **Status:** complete; implemented, verified, and pushed to origin/main
4
+ **Priority:** P2
5
+ **Incident date:** 2026-09-04 PDT / 2026-09-05 UTC
6
+ **Program:** [September 4 Telegram long-haul repairs](TELEGRAM-LONGHAUL-2026-09-04.md)
7
+
8
+ ## Observed failure
9
+
10
+ Current transcription and shell-timeout recovery cards were retired at 19:58 PDT with supersededBy='failure-expired:turn-19', yet unresolved-failure diagnostics still appeared at 20:35. This is contradictory active recovery context; it was not proven to be a terminal blocker.
11
+
12
+ ## Evidence
13
+
14
+ Paths are relative to /home/roko/Documents/Projects/Adjacent. Live files are
15
+ observational references; sanitized deterministic fixtures must reproduce the
16
+ failure without copying credentials or private conversation content.
17
+
18
+ - `telegram_test/.omnius/workboards/telegram-64ac9937dc7647a8-1788572475259-1/active.json`
19
+
20
+ ## Root cause
21
+
22
+ Failure retirement preserves status for audit and marks `supersededBy`.
23
+ Workboard active diagnostic derivation checked failure card ID/status but
24
+ ignored supersession. The same omission affected compact card counts and
25
+ selection, synthesis blockers and verified evidence, and the diagnostics tool
26
+ view. Cached snapshots could retain the old derived warnings even after a
27
+ runtime upgrade.
28
+
29
+ ## Code locations
30
+
31
+ - `packages/execution/src/tools/workboard.ts — active-card selection, diagnostic derivation, projections`
32
+ - `packages/execution/tests/workboard.test.ts`
33
+ - `packages/orchestrator/tests/workboard-run-continuity.test.ts`
34
+
35
+ ## Repair design
36
+
37
+ - Use consistent active-card lifecycle semantics in derived workboard diagnostics and active projections.
38
+ - Exclude superseded/expired obligations from unresolved counts and actionable warnings while preserving historical cards and diagnostic records where explicitly historical.
39
+ - Do not relabel expiration as successful verification.
40
+ - Ensure a genuinely new recurrence creates/activates the correct obligation without resurrecting retired state accidentally.
41
+
42
+ ## Acceptance checklist
43
+
44
+ - [x] An expired/superseded in-progress failure no longer emits active unresolved-failure diagnostics.
45
+ - [x] A current unresolved failure still emits its bounded warning.
46
+ - [x] Retirement preserves the original failure and its audit trail without fabricating verification.
47
+ - [x] Fresh recurrence is represented correctly and legacy cards without supersession retain behavior.
48
+ - [x] Focused workboard/lifecycle tests and execution typecheck pass.
49
+
50
+ ## Implementation and verification log
51
+
52
+ - A single active-card predicate now excludes `supersededBy` cards from compact
53
+ counts, card limits, hidden-card counts, synthesis, and newly derived failure
54
+ and dependency diagnostics. The existing ownership-frontier check uses the
55
+ same predicate.
56
+ - Active diagnostic projection filters recorded or cached historical warnings
57
+ for retired cards before model-output limits are applied. Full JSON, raw
58
+ cards, status, evidence, and append-only events preserve the audit history.
59
+ An explicitly returned compact card includes its supersession marker.
60
+ - Retirement supplies no verification evidence. A live card depending on a
61
+ retired but unverified card remains blocked; the dependent obligation must
62
+ be explicitly reconciled. A new recurrence receives its own active warning.
63
+ - Before the repair, regression fixtures reproduced retired-card
64
+ `unresolved_failure_card`/`repeated_worker_failure` warnings and a board
65
+ containing only retired cards failing to report its empty active frontier.
66
+ - Hermetic verification on 2026-09-05 UTC:
67
+ - `pnpm --filter @omnius/execution exec vitest run tests/workboard.test.ts`:
68
+ **37 tests passed**, including cached-snapshot projections, replay/reload,
69
+ preserved failed evidence, fresh recurrence, and dependency safety.
70
+ - `pnpm --filter @omnius/orchestrator exec vitest run tests/workboard-run-continuity.test.ts`:
71
+ **10 tests passed**.
72
+ - `pnpm --filter @omnius/execution typecheck`: **passed**.
73
+ - Repository delivery commit will be recorded by the parent repair task.
74
+
75
+ ## Delivery
76
+
77
+ - Commits: `3a1731ad` on `origin/main`.
78
+ - Clean `pnpm -r build` passed for all 11 workspace packages/apps. Final aggregate suite evidence is in the linked program ledger.
79
+
80
+ ## Publication boundary
81
+
82
+ Repository repair and deterministic verification are this order's completion
83
+ boundary. The user will publish the package and then perform live validation.
84
+ No package publication, installed-package replacement, runtime-state repair,
85
+ or service restart is authorized by this order.
@@ -0,0 +1,65 @@
1
+ # WO-30: Tool execution must come from structured calls
2
+
3
+ **Status:** complete in repository; verified for user publication
4
+ **Priority:** P1
5
+ **Program:** [September 5 tool-quality and live-behavior follow-up](TOOL-QUALITY-2026-09-05.md)
6
+
7
+ ## Observed failure
8
+
9
+ Live descriptive text containing “shell tool interface” triggered two escalating call corrections. Source review additionally found automatic execution of bash/sh/shell/zsh response fences through a direct shellTool.execute path, bypassing normal tool validation, interruption receipts, and mutation accounting.
10
+
11
+ ## Evidence and code locations
12
+
13
+ - Installed runtime: Omnius 1.0.695, active run telegram-64ac9937dc7647a8-1788593533269-1.
14
+ - Observation window: 2026-09-05 00:32–00:41 PDT.
15
+ - Live evidence: sibling telegram_test/.omnius/context-window-dumps/ and context/steering-ledger.jsonl; sanitized reproductions must avoid private conversation bodies.
16
+ - packages/orchestrator/src/agenticRunner.ts; packages/orchestrator/tests/edit-transport-guidance.test.ts
17
+
18
+ ## Root repair
19
+
20
+ Remove prose/fence-to-execution and tool-name keyword coercion. Keep ordinary assistant prose and examples as text. Execute only actual structured tool calls through the existing admission path. Update the advertised protocol.
21
+
22
+ ## Acceptance
23
+
24
+ - [x] A documentation fence executes no shell; descriptive tool names cause no coercion; genuine structured calls still execute and receive ordinary receipts.
25
+ - [x] Reproduce before repair with deterministic tests.
26
+ - [x] Focused regressions, affected typechecks, and workspace build pass.
27
+ - [x] Independent review is resolved and scoped commits delivered to origin/main.
28
+
29
+ ## Implementation and verification ledger
30
+
31
+ Removed the entire response-fence execution path, the keyword narration classifier, its escalation counter and obsolete file-write heuristic. Updated the runtime protocol. Added structured-tool-authority.test.ts with four shell-language fences, descriptive tool discussion, and a genuine structured-call control. Replaced the obsolete source assertion in edit-transport-guidance.test.ts.
32
+
33
+ Baseline replay: five of six behavioral tests failed against 835fda0f (all four fences executed; descriptive text caused a coercive correction); structured execution passed. After repair: nine tests across both files passed. Logs: /tmp/omnius-wo30-before.log and /tmp/omnius-wo30-test.log. Aggregate build/review and final delivery are recorded in the program ledger.
34
+
35
+ ## Runtime boundary
36
+
37
+ Repair the repository. Preserve live Telegram state, installed package, existing publish/ and docs/DISCOVERY.* changes. No live inference, GPU workloads, service restart, package publication, or external messaging. The user owns publication and later live validation.
38
+
39
+ ### Independent review: XML protocol admission
40
+
41
+ Review reproduced a second bypass: parseTextToolCalls promoted XML examples from arbitrary prose into actions, including for native Qwen requests. The parser now requires explicit host opt-in and a complete standalone envelope sequence. Fenced/prose examples remain intact and inert, malformed batches reject as a whole, and each accepted call receives a unique transaction ID. Thirteen model-profile tests pass (/tmp/omnius-wo30-xml.log). Runner provider/stream and explicit JSON text-mode integration is recorded below.
42
+
43
+ ### Runner protocol integration
44
+
45
+ Provider adapters now normalize only actual wire `tool_calls`. Canonical runner dispatch selects XML from the host's Hermes profile, or JSON from explicit text mode/the individual tools-unsupported retry. Whole standalone envelopes enter ordinary argument validation, tool accounting, interruption, steering, and completion handling. Native transactions take precedence. Code fences, prose, malformed sequences, and structured-output contracts do not grant text-call authority.
46
+
47
+ Removed the retry's fenced-JSON extraction and direct execution branch. Restored the constructor's omitted `textToolMode` option; explicit text mode advertises the standalone JSON grammar and keeps API tools disabled after exposure refreshes. Streaming uses the same admission rules and preserves inert XML examples. A chunk-ending newline no longer becomes an empty internal-marker prefix, preserving fenced output across arbitrary chunk widths.
48
+
49
+ **110 tests passed across seven suites**, including 38 production/provider authority cases, 34 typed-output tests, 22 evidence-steering tests (primary/brute and crash-point recovery), 13 model-profile tests, and three guidance tests. Orchestrator typecheck passed. Logs: `/tmp/omnius-wo30-final.log` and `/tmp/omnius-wo30-types-final.log`. The final explicit-mode exposure-refresh assertion is checked separately in `/tmp/omnius-wo30-json-final.log`. All tests use synthetic workspaces and mocked transports; no live inference or Telegram requests.
50
+
51
+ Final independent review confirmed whole-payload admission and native-call precedence, and found that text catalogs omitted argument schemas. Both text surfaces now serialize host-prepared definitions with full public parameters and a final profile filter. The updated **40 authority cases passed**, including allowed-schema/denied-tool assertions for both explicit mode and the unsupported-tools retry; final orchestrator typecheck passed. Log: `/tmp/omnius-wo30-schema-final.log`. Review findings are resolved; parent owns aggregate clean build and delivery.
52
+
53
+ Aggregate follow-up: keep hidden think blocks out of the visible provider answer while leaving XML tool examples inert. Removing content-derived invocation had also removed the older incidental reasoning strip; the provider boundary now applies only the dedicated reasoning-strip helper. All 67 tests across provider pool/recovery and the 40-case structured-tool authority matrix passed; /tmp/omnius-wo30-think-recovery.log.
54
+
55
+ ## Repository closure — September 5
56
+
57
+ All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
58
+
59
+ Scoped repair commits delivered to origin/main: d2cbec82, e0da0aa6, 4465c9b2, 49f34183. Publication and subsequent live acceptance remain with the user.
60
+
61
+ Source and integration locations:
62
+
63
+ - [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
64
+ - [packages/orchestrator/src/modelProfile.ts](../../../packages/orchestrator/src/modelProfile.ts)
65
+ - [packages/orchestrator/src/typed-model-output.ts](../../../packages/orchestrator/src/typed-model-output.ts)
@@ -0,0 +1,62 @@
1
+ # WO-31: Preserve complete typed runtime policy
2
+
3
+ **Status:** complete in repository; verified for user publication
4
+ **Priority:** P1
5
+ **Program:** [September 5 tool-quality and live-behavior follow-up](TOOL-QUALITY-2026-09-05.md)
6
+
7
+ ## Observed failure
8
+
9
+ The published main prompt is cut at exactly 8,000 characters, mid-sentence. The later medium-model controller and appended agent operating contract never reach the model, despite roughly 80% free context. The separate core reference contract survives.
10
+
11
+ ## Evidence and code locations
12
+
13
+ - Installed runtime: Omnius 1.0.695, active run telegram-64ac9937dc7647a8-1788593533269-1.
14
+ - Observation window: 2026-09-05 00:32–00:41 PDT.
15
+ - Live evidence: sibling telegram_test/.omnius/context-window-dumps/ and context/steering-ledger.jsonl; sanitized reproductions must avoid private conversation bodies.
16
+ - packages/orchestrator/src/agenticRunner.ts; packages/orchestrator/src/context-compiler.ts; packages/orchestrator/src/artifactContract.ts
17
+
18
+ ## Root repair
19
+
20
+ Preserve runtime-owned policy bodies as complete bounded structural sections through both compiler modes and final projection. Continue bounding untrusted/dynamic system material. Use actual request budgeting; do not silently sever invariant tool instructions.
21
+
22
+ ## Acceptance
23
+
24
+ - [x] Large invariant policy reaches final requests intact with headroom; untrusted dynamic blocks remain bounded; tight budgets return an explicit capacity outcome; existing small/medium/large contract tests pass.
25
+ - [x] Reproduce before repair with deterministic tests.
26
+ - [x] Focused regressions, affected typechecks, and workspace build pass.
27
+ - [x] Independent review is resolved and scoped commits delivered to origin/main.
28
+
29
+ ## Implementation and verification ledger
30
+
31
+ Runtime-owned policyScope metadata now takes priority over textual controller/receipt examples. Policy bodies split into complete sections (at most 8,000 characters) without losing text. Legacy compaction reserves their full capacity, emits an explicit error if policy alone exceeds its budget, and neither GC nor the signal distiller can trim/drop their content. Active compilation retains final canonical request-budget admission, including protected-overflow refusal. Dynamic project state remains independently bounded.
32
+
33
+ Added runtime-policy-delivery.test.ts: exact section reconstruction, both compiler modes, actual medium-tier backend requests, explicit impossible capacity, and untrusted textual-marker control. Task-replacement production tests also require the full operating contract after retirement. Updated controller selectors to recognize actual envelope lines rather than quoted examples in the newly visible policy.
34
+
35
+ Baseline replay of the actual active backend request test failed against d2cbec82 (operating policy missing); fixed focused set passes 33 distinct tests across five files. Orchestrator typecheck passes. Logs: /tmp/omnius-wo31-before.log, /tmp/omnius-wo31-test.log, /tmp/omnius-wo31-artifact.log, /tmp/omnius-wo31-types.log. Aggregate verification and independent review remain tracked by the program.
36
+
37
+ ### Final admission follow-up
38
+
39
+ Independent aggregate testing exposed another downstream loss boundary in `context-admission.ts`: policy metadata was ignored, a trailing system message stopped latest-tool protection, and the emergency projection clipped policy or discarded tool transactions to manufacture an admissible request. Four deterministic admission regressions failed before this repair.
40
+
41
+ Final admission now preserves every typed runtime policy section and matches the latest complete tool batch by transaction IDs across trailing system messages. Emergency admission removes only discardable history and retains protected authority, full source results and their metadata. When that complete payload cannot fit, it returns an explicit rejection with zero admitted output tokens. It no longer clips the current user request or replaces it with invented recovery instructions.
42
+
43
+ The small-source headroom fixture now reserves space for the complete policy: synthetic resident noise was reduced from 1,000 to 300 repetitions, and both actual outbound admission budgets are asserted at or below the 16,000-token window. The explicit 9 KB full read remains intact. Updated admission, policy-delivery and evidence-branch suites pass all 34 tests, including both compiler modes through the final admission method, impossible policy/full-read capacity, and emergency tool-batch preservation. Orchestrator typecheck and scoped diff checks pass. All backends are mocked; no live inference or runtime changes were used.
44
+
45
+ ## Runtime boundary
46
+
47
+ Repair the repository. Preserve live Telegram state, installed package, existing publish/ and docs/DISCOVERY.* changes. No live inference, GPU workloads, service restart, package publication, or external messaging. The user owns publication and later live validation.
48
+
49
+ Aggregate recovery fixture reconciliation: the prior synthetic learned ceiling left less capacity than the complete policy itself, so the repaired boundary correctly rejected it. The successful-recovery fixture now supplies a ceiling that can fit the complete policy and asserts byte-for-byte policy preservation, changed request identity, reduced output/prefix and total request fit. Impossible capacities retain separate rejection coverage. All 25 context recovery/admission/runtime-policy tests passed; /tmp/omnius-wo31-recovery.log.
50
+
51
+ ## Repository closure — September 5
52
+
53
+ All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
54
+
55
+ Scoped repair commits delivered to origin/main: 8f3134f0, 08577d79, c6ec5040. Publication and subsequent live acceptance remain with the user.
56
+
57
+ Source and integration locations:
58
+
59
+ - [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
60
+ - [packages/orchestrator/src/context-compiler.ts](../../../packages/orchestrator/src/context-compiler.ts)
61
+ - [packages/orchestrator/src/context-admission.ts](../../../packages/orchestrator/src/context-admission.ts)
62
+ - [packages/orchestrator/src/artifactContract.ts](../../../packages/orchestrator/src/artifactContract.ts)
@@ -0,0 +1,55 @@
1
+ # WO-32: Resolve task scope after referenced evidence arrives
2
+
3
+ **Status:** complete in repository; verified for user publication
4
+ **Priority:** P1
5
+ **Program:** [September 5 tool-quality and live-behavior follow-up](TOOL-QUALITY-2026-09-05.md)
6
+
7
+ ## Observed failure
8
+
9
+ The latest explicit switch request was reconciled before its referenced audio was transcribed. Transcription succeeded and subsequent reads switched to boutique-agent-services, but taskEpoch stayed 1, no boundary archive was created, the canonical goal remained the old task projection, and old raw source stayed active. Persisted lifecycle entries omit parsed reconciliation fields, so omitted disposition versus explicit continue cannot be distinguished.
10
+
11
+ ## Evidence and code locations
12
+
13
+ - Installed runtime: Omnius 1.0.695, active run telegram-64ac9937dc7647a8-1788593533269-1.
14
+ - Observation window: 2026-09-05 00:32–00:41 PDT.
15
+ - Live evidence: sibling telegram_test/.omnius/context-window-dumps/ and context/steering-ledger.jsonl; sanitized reproductions must avoid private conversation bodies.
16
+ - packages/orchestrator/src/{agenticRunner,steeringIntake,typed-model-output}.ts; task scope production tests
17
+
18
+ ## Root repair
19
+
20
+ Make scope decisions explicit and auditable, supporting a bounded evidence-read phase before final reconciliation when the new objective depends on referenced media/source. Preserve the current user request and relevant tool result; prevent unrelated old-plan mutations until scope is resolved. No lexical replacement classifier.
21
+
22
+ ## Acceptance
23
+
24
+ - [x] A replacement referring to untranscribed audio can read that evidence and then retire old scope exactly once; additive/status requests remain continuation; reconciliation decisions are persisted and recovery-safe.
25
+ - [x] Reproduce before repair with deterministic tests.
26
+ - [x] Focused regressions, affected typechecks, and workspace build pass.
27
+ - [x] Independent review is resolved and scoped commits delivered to origin/main.
28
+
29
+ ## Implementation and verification ledger
30
+
31
+ - Added `resolve_after_evidence` with 1–4 exact registered read-call declarations per phase, at most 3 phases. Names, IDs, raw arguments, parsed read modes, profiles, and single-use tickets constrain admission. Unrelated effects and completion stay gated. The canonical `transcribe_file` exception admits only local-file extraction parameters; general tool metadata remains conservative.
32
+ - Applied reconciliations require explicit disposition. The protected user-authority slot and completion obligation remain unresolved through evidence reads. Final replacement adopts the declared evidence transactions and uses the existing durable scope boundary exactly once; explicit additive continuation keeps the epoch.
33
+ - Parsed decisions are persisted with run/epoch/input correlation for both primary and related inputs. A flushed artifact plus atomic selector stores exact authority, goal, evidence tickets, and transcript even before the first scope boundary.
34
+ - Ticket claims are durable before dispatch. Recovery validates identities and normalized transaction arguments, reports interrupted reads without replay, restores the original goal, and preserves the admission gate. Persistence failures cannot grant in-memory continuation.
35
+ - WO35 integration supplies the full scope to file-exploration notes and rebinds tool state at initial run admission and task epoch changes.
36
+ - **77 tests passed** across steering, typed-output, and scope suites before final durability tightening. **22 tests passed** afterward, including primary/brute evidence→replacement and real mocked `runner.run()` crash-point recovery. Orchestrator `tsc --noEmit` passed.
37
+ - Independent review found two defects in the draft: claim persistence occurred too late, and restored receipt anchors lacked argument matching. Both were repaired and covered by the durability tests. Final parent review confirmed argument matching, epoch-before-checkpoint ordering, fail-closed persistence, and exact ticket admission; no additional concrete defect remained in those paths. No live inference or Telegram requests were used.
38
+ - Parent owns final workspace clean build, aggregate checks, and push.
39
+
40
+ ## Runtime boundary
41
+
42
+ Repair the repository. Preserve live Telegram state, installed package, existing publish/ and docs/DISCOVERY.* changes. No live inference, GPU workloads, service restart, package publication, or external messaging. The user owns publication and later live validation.
43
+
44
+ ## Repository closure — September 5
45
+
46
+ All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
47
+
48
+ Scoped repair commits delivered to origin/main: 5276b844. Publication and subsequent live acceptance remain with the user.
49
+
50
+ Source and integration locations:
51
+
52
+ - [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
53
+ - [packages/orchestrator/src/steeringIntake.ts](../../../packages/orchestrator/src/steeringIntake.ts)
54
+ - [packages/orchestrator/src/steeringEvidence.ts](../../../packages/orchestrator/src/steeringEvidence.ts)
55
+ - [packages/orchestrator/src/typed-model-output.ts](../../../packages/orchestrator/src/typed-model-output.ts)
@@ -0,0 +1,55 @@
1
+ # WO-33: Truthful file mutation transactions
2
+
3
+ **Status:** complete in repository; verified for user publication
4
+ **Priority:** P1
5
+ **Scope:** file_write, file_edit, file_patch, batch_edit, notebook_edit, structured_file
6
+
7
+ ## Reproduced failures
8
+
9
+ - A second batch destination with denied write permission caused an exception after the first file changed, without a mutation receipt.
10
+ - Two file_edit calls using the same starting hash both succeeded and lost one edit in 8/8 isolated trials.
11
+ - A target and its symlink alias were staged separately; the second write erased the first edit.
12
+ - file_write continued after an existing-file read error and overwrote a writable, unreadable file without overwrite or hash authorization.
13
+ - Invalid supplied hashes were interpreted as absent. Missing old anchors plus replacement text elsewhere falsely produced already-applied success.
14
+
15
+ ## Implementation
16
+
17
+ The shared file-mutation boundary resolves symlink targets and serializes cooperating tool instances by canonical path. Every destination is read and validated before committing. Complete UTF-8 bodies and original-content rollback copies are staged beside each destination. Existing files are replaced by rename; new files use exclusive linking of their complete staged body. Hashes and aliases are checked again before each commit.
18
+
19
+ If a later commit fails, previous writes are rolled back when their current bytes still match this transaction. The returned result reports partial mutation and the exact unrestored paths when rollback fails; original-content recovery files are retained and named in the error. Invalid UTF-8, unreadable preimages, invalid hash arguments and unsupported hard links fail closed. All six mutation tools use the same boundary and strict hash guard.
20
+
21
+ An absent old anchor now requires fresh inspection even if replacement text occurs elsewhere. True content-preserving edits still return no-op, and cancelling batch edits no longer emit APPLIED lines.
22
+
23
+ ## Guarantees and limits
24
+
25
+ - Serialization covers cooperating calls in this Node process, including different tool instances and symlink aliases. It does not lock independent processes or arbitrary external writers.
26
+ - Individual replacements are atomic. A multi-file batch is staged and rollback-capable, **not crash-atomic**. The process can stop between renames; there is no durable transaction recovery journal.
27
+ - A final pre-commit check reduces external-writer races but cannot provide an operating-system compare-and-swap against unrelated writers.
28
+ - Multiply linked existing files are rejected instead of silently breaking hard-link semantics. Parent directories created for a new file can remain after a failed transaction.
29
+ - Ordinary permission bits are explicitly restored after staging so umask cannot silently narrow an existing mode. Special set-id/sticky bits are rejected. Atomic replacement does not promise preservation of external ACLs, extended attributes or original file ownership.
30
+ - A supplied hash requires an existing preimage; file_write and structured_file cannot recreate a deleted file under its stale hash, including during dry runs.
31
+
32
+ ## Verification
33
+
34
+ Hermetic tests exercise concurrent tools, symlink aliases, staging failure, second-rename failure, rollback failure with retained recovery bytes, read-denied overwrite refusal, exclusive creation races, invalid hashes across all six tools, invalid UTF-8 and hard links, and cancelling no-ops.
35
+
36
+ - Final focused run: **112 tests passed** across file-mutation, file-edit/write/patch, batch-edit, notebook-edit and structured-file suites.
37
+ - `pnpm --filter @omnius/execution typecheck`: **passed**.
38
+ - No live inference, service mutation, installed runtime changes or package publication.
39
+
40
+ ## Repository closure — September 5
41
+
42
+ All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
43
+
44
+ Scoped repair commits delivered to origin/main: a12d6bd2, be0b126b. Publication and subsequent live acceptance remain with the user.
45
+
46
+ Source and integration locations:
47
+
48
+ - [packages/execution/src/tools/file-mutation.ts](../../../packages/execution/src/tools/file-mutation.ts)
49
+ - [packages/execution/src/tools/edit-metadata.ts](../../../packages/execution/src/tools/edit-metadata.ts)
50
+ - [packages/execution/src/tools/file-write.ts](../../../packages/execution/src/tools/file-write.ts)
51
+ - [packages/execution/src/tools/file-edit.ts](../../../packages/execution/src/tools/file-edit.ts)
52
+ - [packages/execution/src/tools/file-patch.ts](../../../packages/execution/src/tools/file-patch.ts)
53
+ - [packages/execution/src/tools/batch-edit.ts](../../../packages/execution/src/tools/batch-edit.ts)
54
+ - [packages/execution/src/tools/notebook-edit.ts](../../../packages/execution/src/tools/notebook-edit.ts)
55
+ - [packages/execution/src/tools/structured-file.ts](../../../packages/execution/src/tools/structured-file.ts)
@@ -0,0 +1,46 @@
1
+ # WO-34: Shell authority and trustworthy process outcomes
2
+
3
+ **Status:** complete in repository; verified for user publication
4
+ **Priority:** P1
5
+
6
+ ## Reproduced failures
7
+
8
+ `find . -delete`, `git branch -D topic` and a `pwd; python3 ...` writer were classified as read-only/concurrency-safe by command-prefix matching. The runner converts that classification into read-only interruption effect tickets.
9
+
10
+ Printing `exit_code: 0` and then exiting 7 produced a failed tool result with structured `primaryExitCode: 0` and `runnerExitCode: 0`. Those fields were reconstructed from rendered command/stdout text. Displayed SIGPIPE markers could similarly override an unrelated real process failure.
11
+
12
+ The legacy process registry marked recovered PID-only sessions killed without sending any signal. Its remote adapter ignored cwd and attempted to obtain exit status with `wait` from a different shell; its stated wait clamp was not applied.
13
+
14
+ ## Repair
15
+
16
+ - Read authority requires one literal command with known read-only semantics. Compound syntax, expansion, interpreter programs, executable rg/find options, unknown options and relevant shell/preprocessor configuration stay effectful. This affects classification, not permission to execute the command.
17
+ - Raw exit code, cwd and timeout state come from spawn callbacks. Exact trailing `echo EXIT=$?` / `EXIT_CODE=$?` forms capture the primary status into the wrapper's private file before emitting display output. Literal single-quoted reporters are ordinary output. Truncation and quoted receipt fields cannot alter primary status.
18
+ - Policy outcomes remain distinct from actual process outcomes. SIGPIPE preview treatment requires an actual observed 141. Tool routing and recovery paths do not invent successful process exits. Elevation uses the actual current cwd.
19
+ - The unused legacy registry is explicitly deprecated toward process-lifecycle/process-async. The environment adapter refuses launch because its API cannot establish owned process and exit receipts. Recovered or remote PID-only sessions cannot be terminated without an owned handle. Owned signals return termination_requested; only observed process exit changes terminal state. Waits now honor their configured clamp.
20
+
21
+ ## Scope and limits
22
+
23
+ The classifier deliberately treats git and interpreter commands as effectful because hooks/configuration/code can execute additional commands. It is not a general shell parser and does not claim unknown commands are safe. The private trailing-reporter capture is implemented for the existing POSIX reporter form; unavailable primary observations remain unknown. Elevated exact reporter forms cannot establish the primary subcommand's status and do not receive a fabricated primary success.
24
+
25
+ Repository searches found no production source imports or root public export of tools/process-registry; its local inspection methods remain for compatibility/tests. Its unsupported environment launch path now fails explicitly instead of pretending to supervise remote work.
26
+
27
+ ## Verification
28
+
29
+ - Combined shell/classifier/receipt/soft-failure/registry run: **115 tests passed across five suites**.
30
+ - Final classifier expansion: **34 tests passed**, including glob expansion and literal double-quote escaping.
31
+ - Execution package typecheck passed.
32
+ - Regression coverage includes real temporary-file deletion classified effectful, receipt-field collisions, false SIGPIPE, private reporter status, literal reporters, routing without execution, recovered termination refusal, termination request versus observed exit, disabled remote launch and fake-timer wait clamping.
33
+ - Reporter capture beyond the stdout cap passed in the combined run.
34
+ - No live inference, installed runtime changes, service changes or publication.
35
+
36
+ ## Repository closure — September 5
37
+
38
+ All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
39
+
40
+ Scoped repair commits delivered to origin/main: a6b6e149. Publication and subsequent live acceptance remain with the user.
41
+
42
+ Source and integration locations:
43
+
44
+ - [packages/execution/src/tools/shell-authority.ts](../../../packages/execution/src/tools/shell-authority.ts)
45
+ - [packages/execution/src/tools/shell.ts](../../../packages/execution/src/tools/shell.ts)
46
+ - [packages/execution/src/tools/process-registry.ts](../../../packages/execution/src/tools/process-registry.ts)
@@ -0,0 +1,56 @@
1
+ # WO-35: Search truth, exploration isolation and source coverage
2
+
3
+ **Status:** complete in repository; verified for user publication
4
+ **Priority:** P1
5
+
6
+ ## Reproduced failures
7
+
8
+ - grep_search and file_explore appended regex text as positional command arguments. Searching for `--version` returned the utility version as successful search evidence.
9
+ - file_explore converted invalid regexes, unavailable paths and failed fallback commands into successful no-match results. It also silently truncated search context.
10
+ - Module-global exploration notes crossed tool instances, working directories and tasks; the runner's unscoped compaction getter could import unrelated findings.
11
+ - file_read admitted fractional ranges and emitted inverted canonical ranges beyond EOF. Requested ends could exceed the actual source while the body was shorter.
12
+ - list_directory omitted entries after 100 without marking incomplete coverage or providing a continuation offset. Failed metadata reads became fabricated zero-byte sizes.
13
+ - find_files advertised path globs, but GNU find's basename-only `-name` treated `**/*.ts` as an impossible basename. During repair, a hermetic fixture also confirmed native Node glob silently treats an unreadable subtree as an empty match set.
14
+
15
+ ## Repair
16
+
17
+ 1. File exploration delegates text search to GrepSearchTool. Both rg and grep delimit pattern data with `-e` and paths with `--`. Only an unavailable rg executable permits fallback. Invalid regexes, inaccessible paths, stderr overflow and process failures remain failures. Context counts are validated. Configured rg preprocessors are disabled. Explicit large files are no longer silently excluded at 2 MiB; timeout and output caps remain bounded. Partial stdout is marked incomplete, and timeout classification uses process fields rather than pattern text in an error message.
18
+ 2. Exploration notes are keyed by canonical working directory and host-bound session, task epoch and owner. Unbound instances have private notes; unscoped export calls return no notes and cannot clear other tasks. Snapshots are defensive copies. A read already in flight keeps its original note collection when the tool is rebound. The latest 128 notes are retained per active scope, with weak registry references allowing abandoned scopes to be collected.
19
+ 3. Canonical reads and exploration chunks validate integer ranges before reading; beyond-EOF requests fail without canonical receipts or saved findings. Headers and receipts report the actual selected end. UTF-8 decoding is strict and preserves BOM bytes in source hashes.
20
+ 4. Directory inventories sort visible entries and accept validated offset/limit pagination. Capped, selected or grouped inventories carry partial materialization metadata and continuation information. Links are identified as links; missing metadata remains unknown.
21
+ 5. File discovery uses strict directory reads with minimatch for actual path/globstar semantics. Unreadable directories fail instead of becoming negative evidence. Directory symlinks and named runtime/dependency directories are skipped. Match, traversal-entry and elapsed-time limits are explicit. A direct minimatch dependency reuses the existing locked 10.2.5 package; installation succeeded offline with a frozen lockfile.
22
+ 6. The tool discovery catalog now describes batch edits as staged writes with rollback receipts, removing its stronger atomicity claim.
23
+
24
+ ## Runner API
25
+
26
+ `getExploreNotes({ workingDir, sessionId, taskEpoch, ownerId })` and `clearExploreNotes` require the same identity used by `FileExploreTool.bindExecutionScope`. The runner must rebind registered tools on task-epoch transitions. Existing no-argument callers are inert for conservative compatibility. Parent integration handles this consumer and epoch binding; this workorder changes no runner code.
27
+
28
+ ## Scope and limits
29
+
30
+ Searches and directory pagination are observations of a live filesystem, not a filesystem snapshot. External changes can move entries between pages. Traversal limits fail with incomplete coverage rather than asserting no files. Search excludes the declared generated/runtime directories, and grep fallback uses extended regex semantics rather than claiming full ripgrep compatibility. File discovery applies full relative-path globs while filename-only patterns match recursively. Working notes are ephemeral, bounded navigation state; normal tool history remains the durable record.
31
+
32
+ Multi-file mutation crash atomicity, cross-process CAS and metadata limits are documented in WO-33. This work does not strengthen those guarantees through discovery text.
33
+
34
+ ## Verification
35
+
36
+ - Five focused suites: **115 tests passed** across search, exploration, glob discovery, canonical file reads and directory inventory.
37
+ - Final timeout-text regression: **24 grep tests passed**.
38
+ - Execution package typecheck and scoped diff checks passed.
39
+ - Behavioral fixtures cover literal option-like searches on rg and grep, a disabled executable rg preprocessor, stderr overflow, large-file search, invalid regexes and paths, partial search coverage, concurrent scope rebinding, defensive note snapshots, fractional and beyond-EOF ranges, exact BOM hashes, paginated coverage without duplicates, symlink inventory, globstar/path matching, unreadable subtrees and capped discovery.
40
+ - No inference, network downloads, live runtime changes, service changes or publication.
41
+
42
+ ## Repository closure — September 5
43
+
44
+ All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
45
+
46
+ Scoped repair commits delivered to origin/main: 8be70469, 5276b844. Publication and subsequent live acceptance remain with the user.
47
+
48
+ Source and integration locations:
49
+
50
+ - [packages/execution/src/tools/grep-search.ts](../../../packages/execution/src/tools/grep-search.ts)
51
+ - [packages/execution/src/tools/glob-find.ts](../../../packages/execution/src/tools/glob-find.ts)
52
+ - [packages/execution/src/tools/file-explore.ts](../../../packages/execution/src/tools/file-explore.ts)
53
+ - [packages/execution/src/tools/explore-tools.ts](../../../packages/execution/src/tools/explore-tools.ts)
54
+ - [packages/execution/src/tools/file-read.ts](../../../packages/execution/src/tools/file-read.ts)
55
+ - [packages/execution/src/tools/list-directory.ts](../../../packages/execution/src/tools/list-directory.ts)
56
+ - [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
@@ -0,0 +1,44 @@
1
+ # WO-36: Preserve transcription evidence and requested capabilities
2
+
3
+ **Status:** complete in repository; verified for user publication
4
+ **Priority:** P1
5
+ **Date:** 2026-09-05
6
+
7
+ ## Reproduced failures
8
+
9
+ - Different recordings with the same basename transcribed within one second
10
+ overwrite the same artifacts; an earlier receipt then reads the later text.
11
+ - A backend error object becomes successful `no_speech` evidence. A valid
12
+ segments-only response displays speech while returning empty model content
13
+ and `no_speech` status.
14
+ - The managed Whisper branch ignores a requested `diarize: true` option.
15
+
16
+ ## Repair and acceptance
17
+
18
+ - Give each accepted transcription its own artifacts and preserve earlier receipts.
19
+ - Validate backend result shape before persisting or classifying silence; derive
20
+ transcript text from valid segments when needed.
21
+ - Reject unsupported managed diarization before model admission, and do not
22
+ silently use that backend as a diarization fallback.
23
+ - Preserve argv literally when invoking the CLI fallback.
24
+ - Exercise all boundaries with mocked backends and temporary files only.
25
+
26
+ ## Verification
27
+
28
+ - `pnpm exec vitest run tests/transcribe-tool-artifacts.test.ts tests/transcribe-python-runtime.test.ts`
29
+ from `packages/execution`: **18 passed**.
30
+ - `pnpm exec tsc --noEmit` from `packages/execution`: passed.
31
+ - The new artifact cases freeze time and verify that each receipt still reads
32
+ its own original text. Failed/malformed results create no evidence artifacts;
33
+ segments-only results preserve speech in model content and storage.
34
+ - No live inference, package installation, service changes, or publication.
35
+
36
+ ## Repository closure — September 5
37
+
38
+ All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
39
+
40
+ Scoped repair commits delivered to origin/main: 8d4e5ce5. Publication and subsequent live acceptance remain with the user.
41
+
42
+ Source and integration locations:
43
+
44
+ - [packages/execution/src/tools/transcribe-tool.ts](../../../packages/execution/src/tools/transcribe-tool.ts)