omnius 1.0.694 → 1.0.696
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/api/py-embed.js +4215 -2598
- package/dist/index.js +23327 -18826
- package/dist/library.js +4564 -2912
- package/dist/python-cuda-runtime.js +512 -0
- package/dist/update-worker.js +4250 -2633
- package/docs/DISCOVERY.json +809 -1
- package/docs/DISCOVERY.md +22 -1
- package/docs/work-orders/runtime-health-remediation/TELEGRAM-LONGHAUL-2026-09-04.md +92 -0
- package/docs/work-orders/runtime-health-remediation/TOOL-QUALITY-2026-09-05.md +84 -0
- package/docs/work-orders/runtime-health-remediation/TRACKER.md +38 -1
- package/docs/work-orders/runtime-health-remediation/WO-25-memory-compiler-eligibility-and-audit.md +81 -0
- package/docs/work-orders/runtime-health-remediation/WO-26-admin-telegram-media-aliases.md +93 -0
- package/docs/work-orders/runtime-health-remediation/WO-27-task-verification-authority.md +140 -0
- package/docs/work-orders/runtime-health-remediation/WO-28-task-scope-evidence-retirement.md +82 -0
- package/docs/work-orders/runtime-health-remediation/WO-29-retired-failure-projections.md +85 -0
- package/docs/work-orders/runtime-health-remediation/WO-30-structured-tool-invocation.md +65 -0
- package/docs/work-orders/runtime-health-remediation/WO-31-complete-runtime-policy.md +62 -0
- package/docs/work-orders/runtime-health-remediation/WO-32-evidence-dependent-steering.md +55 -0
- package/docs/work-orders/runtime-health-remediation/WO-33-file-mutation-transactions.md +55 -0
- package/docs/work-orders/runtime-health-remediation/WO-34-shell-authority-and-results.md +46 -0
- package/docs/work-orders/runtime-health-remediation/WO-35-search-and-exploration-isolation.md +56 -0
- package/docs/work-orders/runtime-health-remediation/WO-36-media-evidence-integrity.md +44 -0
- package/docs/work-orders/runtime-health-remediation/WO-37-browser-and-process-lifecycle.md +86 -0
- package/docs/work-orders/runtime-health-remediation/WO-37-point-localization-cancellation.md +24 -0
- package/docs/work-orders/runtime-health-remediation/WO-38-tool-contract-preservation.md +83 -0
- package/docs/work-orders/runtime-health-remediation/WO-39-telegram-working-indicator.md +76 -0
- package/docs/work-orders/runtime-health-remediation/WO-40-web-content-and-crawl-receipts.md +44 -0
- package/docs/work-orders/runtime-health-remediation/WO-41-completion-evidence-consistency.md +45 -0
- package/docs/work-orders/runtime-health-remediation/WO-42-media-execution-and-configuration.md +79 -0
- package/npm-shrinkwrap.json +5 -5
- package/package.json +1 -1
package/docs/DISCOVERY.md
CHANGED
|
@@ -532,8 +532,10 @@ Daemon equivalents are `GET /v1/discovery/bootstrap`, `GET /v1/discovery?q=<inte
|
|
|
532
532
|
| `guide.work-orders-runtime-health-remediation-h6-main-tui-wo-05-confirmation-evidence-uppercase` | H6 Main-TUI and WO-05 Confirmation Evidence | This closure wires the top-level TUI FullSubAgentTool to live parent-context headroom and the exact owning task cancellation signal. It also confirms the WO-05 completion authority after the interruption, safe-media, workboard, and canonical-context changes. |
|
|
533
533
|
| `guide.work-orders-runtime-health-remediation-live-telegram-log-review-2026-09-03-uppercase` | Live Telegram Log Review, 2026-09-03 | This review is read-only. It examines the persisted Telegram conversation and intake lifecycle records in /home/roko/Documents/Projects/Adjacent/telegramtest/.omnius. The running global Omnius package reports version 1.0.687. This means the observations describe the deployed pre-publication runtime, not the newer source commits on this repository's main bran |
|
|
534
534
|
| `guide.work-orders-runtime-health-remediation-readme-uppercase` | Runtime Health Remediation Program | Program ID: RHR-2026-09-02 Status: active, implementation authorized Owner: Omnius runtime, orchestration, execution, memory, and Telegram packages External dependency: /home/roko/Desktop/ollama-unify Last reconciled: 2026-09-03 |
|
|
535
|
+
| `guide.work-orders-runtime-health-remediation-telegram-longhaul-2026-09-04-uppercase` | September 4 Telegram long-haul remediation | Status: complete; all five root repairs verified and delivered to origin/main Baseline: Omnius main 69c4a160; installed package 1.0.694 Incident run: telegram-64ac9937dc7647a8-1788572475259-1 Observation window: 2026-09-04 18:41–20:43 PDT Owner: repository repair; user owns publication and subsequent live testing |
|
|
536
|
+
| `guide.work-orders-runtime-health-remediation-tool-quality-2026-09-05-uppercase` | September 5 tool-quality and live-behavior follow-up | Status: WO-30 through WO-42 complete in repository; full verification passed; delivered to origin/main for user publication Repository baseline: 501e8394 on origin/main Verified source head: 7005aa41; subsequent closure changes are documentation only Observed running package: 1.0.695, verified from its actual executable/package path User scope: monitor live |
|
|
535
537
|
| `guide.work-orders-runtime-health-remediation-traceability-uppercase` | Runtime Health Remediation Traceability Matrix | Every completed row links to a deterministic test or operational receipt. A commit hash alone is not proof. A model statement is not proof. Live logs may support a canary only after the hermetic and deterministic gates pass. A status that names pending rollout or live work is intentionally not a completion claim. |
|
|
536
|
-
| `guide.work-orders-runtime-health-remediation-tracker-uppercase` | Runtime Health Remediation Master Tracker | Authority: canonical granular tracker for RHR-2026-09-02 Checked-item rule: code, focused tests, and named evidence must all exist Last reconciled: 2026-09-
|
|
538
|
+
| `guide.work-orders-runtime-health-remediation-tracker-uppercase` | Runtime Health Remediation Master Tracker | Authority: canonical granular tracker for RHR-2026-09-02 Checked-item rule: code, focused tests, and named evidence must all exist Last reconciled: 2026-09-05 |
|
|
537
539
|
| `guide.work-orders-runtime-health-remediation-wo-00-hermetic-test-boundary-uppercase` | WO-00: Hermetic Test Boundary | Status: deterministic acceptance complete Evidence: inference network inventory and audit receipt for commit 8708e449 Risk: high, because current unit tests can contact the live broker Depends on: none |
|
|
538
540
|
| `guide.work-orders-runtime-health-remediation-wo-00-inference-network-inventory-uppercase` | WO-00 Inference Network Inventory | This document describes the machine-checked boundary for production inference clients. The canonical inventory is test-support/inference-network-registry.ts. |
|
|
539
541
|
| `guide.work-orders-runtime-health-remediation-wo-01-inference-pressure-scheduler-uppercase` | WO-01: Foreground and Background Inference Pressure Scheduler | Status: deterministic acceptance complete Evidence: production Telegram cognition integration and controlled-load foreground exclusion Risk: critical, because background work currently competes with user work Depends on: WO-00 |
|
|
@@ -563,6 +565,25 @@ Daemon equivalents are `GET /v1/discovery/bootstrap`, `GET /v1/discovery?q=<inte
|
|
|
563
565
|
| `guide.work-orders-runtime-health-remediation-wo-22-autocomplete-footer-geometry-uppercase` | WO-22: Autocomplete Footer Geometry Integrity | Scope: CLI TUI footer, DirectInput redraw scheduling, active content writes |
|
|
564
566
|
| `guide.work-orders-runtime-health-remediation-wo-23-long-horizon-control-plane-reconciliation-uppercase` | WO-23: Long-horizon control-plane reconciliation | The restored telegramtest run tui-3869250-mtlvp324-1788468819792-1 continued a large pentest-tooling frontend task across a process restart. It made real source mutations and advanced from the first implementation group to todo-p2a. The run was not dead and its model context had ample nominal capacity. |
|
|
565
567
|
| `guide.work-orders-runtime-health-remediation-wo-24-telegram-logical-request-variant-identity-uppercase` | WO-24: Telegram logical request variant identity | Telegram follow-up requests could finish a streamed inference with no visible contract and then fail their one unary compatibility request with HTTP 409: logicalrequestconflict. The observed stream used all 300 completion tokens and ended with finish=length. |
|
|
568
|
+
| `guide.work-orders-runtime-health-remediation-wo-25-memory-compiler-eligibility-and-audit-uppercase` | WO-25: Memory compiler eligibility and outcome audit | Status: complete; implemented, verified, and pushed to origin/main Priority: P1 Incident date: 2026-09-04 PDT / 2026-09-05 UTC Program: September 4 Telegram long-haul repairs |
|
|
569
|
+
| `guide.work-orders-runtime-health-remediation-wo-26-admin-telegram-media-aliases-uppercase` | WO-26: Admin Telegram media alias resolution | Status: complete; implemented, verified, and pushed to origin/main Priority: P1 Incident date: 2026-09-04 PDT / 2026-09-05 UTC Program: September 4 Telegram long-haul repairs |
|
|
570
|
+
| `guide.work-orders-runtime-health-remediation-wo-27-task-verification-authority-uppercase` | WO-27: Separate mutation observations from task verification | Status: complete; implemented, verified, and pushed to origin/main Priority: P1 Incident date: 2026-09-04 PDT / 2026-09-05 UTC Program: September 4 Telegram long-haul repairs |
|
|
571
|
+
| `guide.work-orders-runtime-health-remediation-wo-28-task-scope-evidence-retirement-uppercase` | WO-28: Retire obsolete task evidence after explicit scope changes | Status: complete; implemented, verified, and pushed to origin/main Priority: P2 Incident date: 2026-09-04 PDT / 2026-09-05 UTC Program: September 4 Telegram long-haul repairs |
|
|
572
|
+
| `guide.work-orders-runtime-health-remediation-wo-29-retired-failure-projections-uppercase` | WO-29: Exclude retired failures from active diagnostics | Status: complete; implemented, verified, and pushed to origin/main Priority: P2 Incident date: 2026-09-04 PDT / 2026-09-05 UTC Program: September 4 Telegram long-haul repairs |
|
|
573
|
+
| `guide.work-orders-runtime-health-remediation-wo-30-structured-tool-invocation-uppercase` | WO-30: Tool execution must come from structured calls | Status: complete in repository; verified for user publication Priority: P1 Program: September 5 tool-quality and live-behavior follow-up |
|
|
574
|
+
| `guide.work-orders-runtime-health-remediation-wo-31-complete-runtime-policy-uppercase` | WO-31: Preserve complete typed runtime policy | Status: complete in repository; verified for user publication Priority: P1 Program: September 5 tool-quality and live-behavior follow-up |
|
|
575
|
+
| `guide.work-orders-runtime-health-remediation-wo-32-evidence-dependent-steering-uppercase` | WO-32: Resolve task scope after referenced evidence arrives | Status: complete in repository; verified for user publication Priority: P1 Program: September 5 tool-quality and live-behavior follow-up |
|
|
576
|
+
| `guide.work-orders-runtime-health-remediation-wo-33-file-mutation-transactions-uppercase` | WO-33: Truthful file mutation transactions | Status: complete in repository; verified for user publication Priority: P1 Scope: filewrite, fileedit, filepatch, batchedit, notebookedit, structuredfile |
|
|
577
|
+
| `guide.work-orders-runtime-health-remediation-wo-34-shell-authority-and-results-uppercase` | WO-34: Shell authority and trustworthy process outcomes | Status: complete in repository; verified for user publication Priority: P1 |
|
|
578
|
+
| `guide.work-orders-runtime-health-remediation-wo-35-search-and-exploration-isolation-uppercase` | WO-35: Search truth, exploration isolation and source coverage | Status: complete in repository; verified for user publication Priority: P1 |
|
|
579
|
+
| `guide.work-orders-runtime-health-remediation-wo-36-media-evidence-integrity-uppercase` | WO-36: Preserve transcription evidence and requested capabilities | Status: complete in repository; verified for user publication Priority: P1 Date: 2026-09-05 |
|
|
580
|
+
| `guide.work-orders-runtime-health-remediation-wo-37-browser-and-process-lifecycle-uppercase` | WO-37: Preserve browser ownership and process outcome truth | Status: complete in repository; verified for user publication Priority: P1 Date: 2026-09-05 |
|
|
581
|
+
| `guide.work-orders-runtime-health-remediation-wo-37-point-localization-cancellation-uppercase` | WO-37 follow-up: Cancel the active point-localization client | Status: complete in repository; verified for user publication Parent: WO-37 browser and process lifecycle |
|
|
582
|
+
| `guide.work-orders-runtime-health-remediation-wo-38-tool-contract-preservation-uppercase` | WO-38: Preserve tool contracts across adapters and policy resolution | Status: complete in repository; verified for user publication Priority: P1 Date: 2026-09-05 PDT |
|
|
583
|
+
| `guide.work-orders-runtime-health-remediation-wo-39-telegram-working-indicator-uppercase` | WO-39: Keep Telegram's working indicator alive through active work | Status: complete in repository; verified for user publication Priority: P1 Date: 2026-09-05 PDT |
|
|
584
|
+
| `guide.work-orders-runtime-health-remediation-wo-40-web-content-and-crawl-receipts-uppercase` | WO-40: Preserve retrieved content and validate crawler responses | Status: complete in repository; verified for user publication Priority: P1 Date: 2026-09-05 |
|
|
585
|
+
| `guide.work-orders-runtime-health-remediation-wo-41-completion-evidence-consistency-uppercase` | WO-41: Expose verification coverage and preserve typed completion authority | Status: complete in repository; verified for user publication Priority: P1 Program: September 5 follow-up |
|
|
586
|
+
| `guide.work-orders-runtime-health-remediation-wo-42-media-execution-and-configuration-uppercase` | WO-42: Media execution and configuration boundaries | Status: complete in repository; verified for user publication Priority: P1 |
|
|
566
587
|
| `guide.work-orders-telegram-dmn-wo-22-dmn-outreach-and-learning-uppercase` | WO-22: DMN outreach, DM sharing, and outcome learning | On 2026-09-03 at 17:07 PDT the bot posted in the OMNIUS group without being addressed. The operator asked whether this was self-induced reflection. |
|
|
567
588
|
| `guide.work-orders-telegram-dropbear-context-rca-workorder` | Telegram Dropbear Context Engineering RCA Work Order | Observed run: /home/roko/Documents/Projects/Adjacent/telegramtest/.omnius, run id 1782873796963-i5r7mv. |
|
|
568
589
|
| `guide.work-orders-wo-am-gaps-uppercase` | Associative Memory Gap Work Orders | Generated: 2026-04-13 Source: Deep audit of multimodal associative memory systems Status: READY FOR IMPLEMENTATION |
|
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
# September 4 Telegram long-haul remediation
|
|
2
|
+
|
|
3
|
+
**Status:** complete; all five root repairs verified and delivered to origin/main
|
|
4
|
+
**Baseline:** Omnius main 69c4a160; installed package 1.0.694
|
|
5
|
+
**Incident run:** telegram-64ac9937dc7647a8-1788572475259-1
|
|
6
|
+
**Observation window:** 2026-09-04 18:41–20:43 PDT
|
|
7
|
+
**Owner:** repository repair; user owns publication and subsequent live testing
|
|
8
|
+
|
|
9
|
+
## Work orders and dependency order
|
|
10
|
+
|
|
11
|
+
| Work order | Finding | Status | Delivery |
|
|
12
|
+
| --- | --- | --- | --- |
|
|
13
|
+
| [WO-25](WO-25-memory-compiler-eligibility-and-audit.md) | Memory compiler eligibility and outcome audit | complete | cb316de6 |
|
|
14
|
+
| [WO-26](WO-26-admin-telegram-media-aliases.md) | Admin Telegram media alias resolution | complete | bc494850 |
|
|
15
|
+
| [WO-27](WO-27-task-verification-authority.md) | Separate mutation observations from task verification | complete | ce40bfca, cf966ea3, dc60e5bf |
|
|
16
|
+
| [WO-28](WO-28-task-scope-evidence-retirement.md) | Retire obsolete task evidence after explicit scope changes | complete | 24321b30 |
|
|
17
|
+
| [WO-29](WO-29-retired-failure-projections.md) | Exclude retired failures from active diagnostics | complete | 3a1731ad |
|
|
18
|
+
|
|
19
|
+
WO-25 comes first on the shared request boundary. WO-26, WO-27, and WO-29
|
|
20
|
+
can be implemented independently. WO-28 integrates task-scope retirement with
|
|
21
|
+
WO-25 after the existing steering boundary is traced. Shared agenticRunner.ts
|
|
22
|
+
changes are serialized by the primary implementer.
|
|
23
|
+
|
|
24
|
+
## Cross-cutting invariants
|
|
25
|
+
|
|
26
|
+
- Repair causes in Omnius; do not edit the active telegram_test workload or logs.
|
|
27
|
+
- Preserve existing dirty publish/ and docs/DISCOVERY.* changes.
|
|
28
|
+
- Retain graph-proven active evidence, user authority, valid tool transactions,
|
|
29
|
+
public Telegram isolation, and current explicit verifier contracts.
|
|
30
|
+
- Add deterministic regressions for the observed failures, including tests that
|
|
31
|
+
assert useful outcomes rather than invocation alone.
|
|
32
|
+
- Use the hermetic network boundary. No live inference, model load, CUDA/service
|
|
33
|
+
changes, package publication, or installed-package replacement.
|
|
34
|
+
- Check work-order boxes only after implementation and named verification pass.
|
|
35
|
+
- Commit each coherent verified repair using only its owned files; fetch and
|
|
36
|
+
push the tracked origin/main branch without rewriting history.
|
|
37
|
+
|
|
38
|
+
## Delivery gates
|
|
39
|
+
|
|
40
|
+
- [x] All five work orders have implemented root repairs and passing regressions.
|
|
41
|
+
- [x] Compiler outcome telemetry survives final request projection.
|
|
42
|
+
- [x] Affected package regression suites and typechecks pass.
|
|
43
|
+
- [x] Clean workspace-package rebuild succeeds (without bundling publish/).
|
|
44
|
+
- [x] Independent review findings are resolved.
|
|
45
|
+
- [x] Work-order links and acceptance evidence are reconciled.
|
|
46
|
+
- [x] Scoped commits are pushed to origin/main.
|
|
47
|
+
- [x] Publication handoff records the final commit, verification, and operator
|
|
48
|
+
canaries for the user to run after publishing.
|
|
49
|
+
|
|
50
|
+
## Verification ledger
|
|
51
|
+
|
|
52
|
+
- WO-25: 46 compiler/audit/canonical regression tests passed; ineligible high-volume output replay made zero compiler calls. Outcome projection and error-cache tests passed.
|
|
53
|
+
- WO-26: 71 media-focused regressions passed, including admin aliases, replies, OCR batch behavior, explicit paths, and public-chat isolation.
|
|
54
|
+
- WO-27: initial 93 completion tests passed, then 168 broader authority/command-trust tests passed. Final parser/observability refinements passed 93 focused tests with actual declared-verifier receipts.
|
|
55
|
+
- WO-29: 37 workboard and 10 continuity tests passed; superseded failures stay in audit history and leave active diagnostics.
|
|
56
|
+
- WO-28: 13 runner scope tests, 11 pure-boundary tests, 25 interruption lifecycle tests, 10 archive-store tests, and two full mocked runner replays passed. The production replays exercise both primary and brute-force loops.
|
|
57
|
+
- Clean rebuild: `pnpm -r clean`, removal of remaining workspace `tsconfig.tsbuildinfo` caches (excluding node_modules and publish), then `pnpm -r build`: **11 workspace projects passed**.
|
|
58
|
+
- Final full execution suite: **1,511 passed, 3 skipped**, 141 test files.
|
|
59
|
+
- Final full CLI suite: **2,480 passed**, 263 test files. This includes the mocked frontend transport lifecycle gate.
|
|
60
|
+
- Final full orchestrator suite: **2,519 passed, 1 skipped**, 208 test files, with `OMNIUS_SQLITE_TESTS=1`.
|
|
61
|
+
- Aggregate result: **6,510 tests passed, 4 skipped**. No test failed in the final full runs. The last directory-sync refinement passed another 15 scope tests and a workspace build.
|
|
62
|
+
- Full suites use `vitest run --maxWorkers=4 --minWorkers=1` through each package and the existing hermetic network setup.
|
|
63
|
+
- Initial aggregate run exposed the core-rule truncation regression, a parser snapshot loaded during active edits, and an obsolete observability fixture. Core rules now retain a separate typed policy slot; parser and fixture corrections passed fresh targeted replays before the final full run.
|
|
64
|
+
- Local work-order links: **11 checked, zero broken**. Pre-existing publish/ and docs/DISCOVERY.* changes were excluded from every task commit.
|
|
65
|
+
|
|
66
|
+
Historic checked claims in TRACKER.md are not acceptance evidence for these newly found defects.
|
|
67
|
+
|
|
68
|
+
## Delivery handoff
|
|
69
|
+
|
|
70
|
+
Code is delivered through `24321b30` on `origin/main`; each work-order table entry records its scoped repair commits. This document and the master tracker close the repository implementation boundary. The existing installed package, running Telegram workload, and publication files were not replaced by this repair.
|
|
71
|
+
|
|
72
|
+
The workspace rebuild from clean compiled outputs passed. The user can now perform the normal publish-package bundling/publication procedure and run the following live canaries. Those canaries remain explicitly unexecuted here; deterministic test success is not represented as live-model validation.
|
|
73
|
+
|
|
74
|
+
## Operator canaries after publication
|
|
75
|
+
|
|
76
|
+
These are deferred to the user, not unfinished implementation steps:
|
|
77
|
+
|
|
78
|
+
1. Send an admin-DM voice message and reference it by reply and message ID;
|
|
79
|
+
confirm immediate scoped resolution and correct transcription attribution.
|
|
80
|
+
2. Resume long work with large tool outputs but ample headroom; confirm zero
|
|
81
|
+
ineligible compiler calls and inspect retained compiler outcome telemetry.
|
|
82
|
+
3. Replace the task with a different project; confirm old raw source leaves the
|
|
83
|
+
active context while new instructions and media evidence remain available.
|
|
84
|
+
4. Build a scaffold, install dependencies, then inspect completion; installation
|
|
85
|
+
alone must not claim behavioral verification. Run the declared checks and
|
|
86
|
+
confirm fresh verifier evidence permits completion.
|
|
87
|
+
5. Let a bounded failure expire or supersede it; confirm active diagnostics
|
|
88
|
+
remove it while audit history remains.
|
|
89
|
+
|
|
90
|
+
Any token-generating canary must first satisfy the repository endpoint/model
|
|
91
|
+
and exact physical GPU preflight gate.
|
|
92
|
+
|
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
# September 5 tool-quality and live-behavior follow-up
|
|
2
|
+
|
|
3
|
+
**Status:** WO-30 through WO-42 complete in repository; full verification passed; delivered to origin/main for user publication
|
|
4
|
+
**Repository baseline:** 501e8394 on origin/main
|
|
5
|
+
**Verified source head:** 7005aa41; subsequent closure changes are documentation only
|
|
6
|
+
**Observed running package:** 1.0.695, verified from its actual executable/package path
|
|
7
|
+
**User scope:** monitor live behavior, audit tool implementations, remedy demonstrated poor practices, keep Telegram typing active, and track durable work orders to completion. Publication belongs to the user.
|
|
8
|
+
|
|
9
|
+
## Live observations
|
|
10
|
+
|
|
11
|
+
The observed run, telegram-64ac9937dc7647a8-1788593533269-1, ran from 00:32:13 through 00:58:09 PDT on September 5. Polling and spool processing remained healthy with zero consecutive polling failures. It completed at task epoch 1 without ambiguous interruption effects.
|
|
12
|
+
|
|
13
|
+
The previous compiler stall did not recur. Raw discovery exceeded 48,000 characters while the request retained substantial context headroom, and no memory compiler request was made. Two early trajectory-grounding calls took about 55 seconds combined. The media alias repair worked: transcribe_file resolved message_id:2755 and returned usable evidence in about 4.4 seconds.
|
|
14
|
+
|
|
15
|
+
User steering redirected work into boutique-agent-services, and subsequent reads and edits followed that target. However, the run retained epoch 1, the old goal/source context, and no scope archive. This exposed the need to defer a scope decision until referenced media has actually been read. The outgoing runtime system prompt was also clipped mid-section despite headroom, and innocent descriptions of tool interfaces triggered corrective feedback.
|
|
16
|
+
|
|
17
|
+
The run changed boutique/src/executor.ts and boutique/src/index.ts, ran checks and probes, and delivered a final report. Its terminal record reported completed/ready while its audit ledger retained verification_missing. Several check commands used trailing status-reporting echoes; a zero wrapper status alone does not establish verifier success. Completion obligations and verification coverage are distinct: explicit required checks gate completion, while observed gaps must remain visible in terminal audit data.
|
|
18
|
+
|
|
19
|
+
## Work orders
|
|
20
|
+
|
|
21
|
+
| Order | Confirmed issue and repair area | Status | Scoped delivery to origin/main |
|
|
22
|
+
| --- | --- | --- | --- |
|
|
23
|
+
| [WO-30](WO-30-structured-tool-invocation.md) | Stop prose/fences/XML examples becoming actions; require a host-selected whole tool protocol | complete | d2cbec82, e0da0aa6, 4465c9b2, 49f34183 |
|
|
24
|
+
| [WO-31](WO-31-complete-runtime-policy.md) | Preserve complete typed runtime policy and current evidence through final admission | complete | 8f3134f0, 08577d79, c6ec5040 |
|
|
25
|
+
| [WO-32](WO-32-evidence-dependent-steering.md) | Durable bounded evidence reads before scope decisions, exact tickets, recovery and epoch retirement | complete | 5276b844 |
|
|
26
|
+
| [WO-33](WO-33-file-mutation-transactions.md) | Guard file replacement, hashes, aliases, concurrency, modes and rollback receipts | complete | a12d6bd2, be0b126b |
|
|
27
|
+
| [WO-34](WO-34-shell-authority-and-results.md) | Conservative shell authority and typed process outcome receipts | complete | a6b6e149 |
|
|
28
|
+
| [WO-35](WO-35-search-and-exploration-isolation.md) | Search options/errors, glob semantics, scoped exploration notes and truthful read/list coverage | complete | 8be70469, 5276b844 |
|
|
29
|
+
| [WO-36](WO-36-media-evidence-integrity.md) | Unique transcript identities, validated backend results and explicit diarization support | complete | 8d4e5ce5 |
|
|
30
|
+
| [WO-37](WO-37-browser-and-process-lifecycle.md) | Session/process ownership, bounded transport and workers, cancellation and startup cleanup | complete | 0d6dcc66, fe3c6487, 52cb2623, 484938c1, 72382afa, 4a072473 |
|
|
31
|
+
| [WO-38](WO-38-tool-contract-preservation.md) | Preserve typed results, parsed inputs, isolated policies and streaming execution wrappers | complete | e4598f71, cae592bc, 48fd2410 |
|
|
32
|
+
| [WO-39](WO-39-telegram-working-indicator.md) | Keep three-second typing active throughout DM/group/topic work and final delivery | complete | 53216239, c17ae247 |
|
|
33
|
+
| [WO-40](WO-40-web-content-and-crawl-receipts.md) | Preserve plain/JSON content, honest cache metadata and validated crawl receipts | complete | 9f96cfcb |
|
|
34
|
+
| [WO-41](WO-41-completion-evidence-consistency.md) | Separate verification audit coverage from explicit completion authority and typed expected effects | complete | 6ee251c3 |
|
|
35
|
+
| [WO-42](WO-42-media-execution-and-configuration.md) | Literal capture arguments, isolated transcription setup, bounded media lifetimes and cancellation | complete | 7bc9106d, 52cb2623, b82f88a9, ed341587, 793c5302 |
|
|
36
|
+
|
|
37
|
+
The [point-localization cancellation follow-up](WO-37-point-localization-cancellation.md) is part of WO-37. All independent review findings are resolved; each order includes concrete source locations, acceptance evidence and limitations.
|
|
38
|
+
|
|
39
|
+
## Review coverage and limits
|
|
40
|
+
|
|
41
|
+
The starting execution tool directory inventory contained 143 source files, including helpers. Deep behavioral review covered filesystem mutation, shell/process management, read/search/exploration, transcription, browser/video/capture, web fetch/crawl, registration/adapters, and runner admission/context/steering. Inventory coverage is distinct from behavioral proof; this is not a claim that every tool and external service is proven correct.
|
|
42
|
+
|
|
43
|
+
Tests use temporary fixtures and mocked inference, Telegram, browser and media transports. This work does not publish, send real Telegram messages, start models, alter services, or repair the monitored workspace in place. The observed live behavior remains evidence for installed 1.0.695; new repository repairs require the user's publication and subsequent field test. A final read-only check at 02:24 PDT still found the same latest completion record, last updated at 00:58:09 PDT; this is not a live acceptance test of the new source.
|
|
44
|
+
|
|
45
|
+
File mutations provide process-local ownership, guarded per-file atomic replacement and explicit rollback diagnostics. They do not claim cross-process compare-and-swap or crash-atomic multi-file transactions. Owned child processes and local browser handles are canceled and drained. Shared Comfy workflows or a Moondream SDK operation can continue after client cancellation where the backend provides no scoped cancellation API; returned diagnostics say so. The implementation does not kill another client's shared service to manufacture a successful Stop receipt.
|
|
46
|
+
|
|
47
|
+
## Delivery gates
|
|
48
|
+
|
|
49
|
+
- [x] Consolidated audit findings have reproductions and work orders.
|
|
50
|
+
- [x] Every confirmed defect in this pass is repaired and independently reviewed.
|
|
51
|
+
- [x] Affected regression suites and clean workspace build pass.
|
|
52
|
+
- [x] Work orders record scope, evidence, results and scoped commits delivered to origin/main.
|
|
53
|
+
- [x] Handoff distinguishes repository verification from unpublished/live behavior.
|
|
54
|
+
- Publication and the live Telegram canary belong to the user after this handoff.
|
|
55
|
+
|
|
56
|
+
## Verification ledger
|
|
57
|
+
|
|
58
|
+
Final verification covers source through 7005aa41. Tests use the repository's hermetic network boundary, mock external systems, and use temporary synthetic workers for process-lifecycle cases.
|
|
59
|
+
|
|
60
|
+
| Verification | Result | Local execution log |
|
|
61
|
+
| --- | --- | --- |
|
|
62
|
+
| Complete orchestrator suite, SQLite tests enabled | 215 suites; 2,624 passed, 1 skipped | /tmp/omnius-tool-quality-orchestrator-verified.log |
|
|
63
|
+
| Complete execution suite after clean build | 160 suites; 1,747 passed, 3 skipped | /tmp/omnius-tool-quality-execution-verified.log |
|
|
64
|
+
| Complete CLI suite after clean build | 266 suites; 2,514 passed | /tmp/omnius-tool-quality-cli-verified.log |
|
|
65
|
+
| Clean all workspace packages, remove residual TypeScript build caches, rebuild | All 11 workspace packages passed | /tmp/omnius-tool-quality-clean-build.log |
|
|
66
|
+
| CUDA preparation helper package entry and artifact policy | Actual source build configuration produced an importable 20,494-byte temporary module, no sourcemap; 16 policy tests passed | /tmp/omnius-cuda-worker-package.log; /tmp/omnius-cuda-worker-package-tests.log |
|
|
67
|
+
|
|
68
|
+
**Total: 641 passing suites, 6,885 passing tests, 4 skipped tests, zero failures.** The full CLI suite includes web-chat-transport-lifecycle, web-ui-client-runtime and web-ui-script. Focused review and before-repair reproductions remain recorded in each work order. Test commands used vitest run --maxWorkers=4 --minWorkers=1 in the respective workspace; the orchestrator run additionally set OMNIUS_SQLITE_TESTS=1. Tools: Node 24.14.0, pnpm 9.15.4, npm 11.9.0.
|
|
69
|
+
|
|
70
|
+
Aggregate review corrected an initial completion overreach: generic mutation alone does not impose mandatory verification; 6ee251c3 preserves audit coverage separately from explicit completion obligations. The full-policy recovery and native-edit fixtures now assert the intended full-policy and exact-anchor contracts. Provider recovery also retains reasoning stripping without interpreting content as tool authority. All affected full suites passed after these corrections.
|
|
71
|
+
|
|
72
|
+
The standalone CUDA preparation helper is now required by both package audits and built by the normal publish script. Verification used temporary output because the user's publish staging directory already had unrelated changes. A complete publication tarball was not rebuilt or published during this pass; the user must run the existing clean-build/bundle/pack publish SOP.
|
|
73
|
+
|
|
74
|
+
## User publication and live acceptance
|
|
75
|
+
|
|
76
|
+
After publishing and running the new package, verify the installed executable's actual version before attributing behavior to these repairs.
|
|
77
|
+
|
|
78
|
+
1. In a DM, start work that includes media preparation or a silent tool interval of at least 18 seconds. Typing should refresh every three seconds through native drafts, silent work and final delivery.
|
|
79
|
+
2. Repeat in public and private groups, including a topic. Steer the active run and confirm its one working heartbeat persists; then Stop and confirm activity retires without a stale timer.
|
|
80
|
+
3. Exercise a normal failure and cancellation during browser/media setup. Owned subprocesses must drain; any shared-server continuation limitation must be explicit.
|
|
81
|
+
4. Switch objectives through referenced audio: evidence is read before the explicit scope decision, old scope retires once, and the selected evidence survives recovery.
|
|
82
|
+
5. Inspect actual outgoing context for complete runtime policy and retained source evidence. A capacity refusal must be explicit, and verification audit gaps must remain separate from configured completion requirements.
|
|
83
|
+
|
|
84
|
+
These are future field checks, not completed live evidence. No installed package, Telegram service or monitored workspace was changed in this repair pass.
|
|
@@ -2,7 +2,44 @@
|
|
|
2
2
|
|
|
3
3
|
**Authority:** canonical granular tracker for RHR-2026-09-02
|
|
4
4
|
**Checked-item rule:** code, focused tests, and named evidence must all exist
|
|
5
|
-
**Last reconciled:** 2026-09-
|
|
5
|
+
**Last reconciled:** 2026-09-05
|
|
6
|
+
|
|
7
|
+
## September 5 tool-quality and live-behavior follow-up
|
|
8
|
+
|
|
9
|
+
Current follow-up authority: [TOOL-QUALITY-2026-09-05.md](TOOL-QUALITY-2026-09-05.md).
|
|
10
|
+
|
|
11
|
+
- [x] WO-30 structured tool invocation and prose safety.
|
|
12
|
+
- [x] WO-31 complete runtime policy delivery.
|
|
13
|
+
- [x] WO-32 evidence-dependent scope reconciliation.
|
|
14
|
+
- [x] Tool-family audit findings reproduced, repaired, and delivered.
|
|
15
|
+
- [x] Aggregate regression/build verification and publication handoff.
|
|
16
|
+
- [x] WO-33: File mutation transactions
|
|
17
|
+
- [x] WO-34: Shell authority and process results
|
|
18
|
+
- [x] WO-35: Search and exploration isolation
|
|
19
|
+
- [x] WO-36: Media evidence integrity
|
|
20
|
+
- [x] WO-37: Browser and process lifecycle
|
|
21
|
+
- [x] WO-38: Tool contract preservation
|
|
22
|
+
- [x] WO-39: Telegram working indicator
|
|
23
|
+
- [x] WO-40: Web content and crawl receipts
|
|
24
|
+
- [x] WO-41: Verification audit and typed completion consistency
|
|
25
|
+
- [x] WO-42: Media execution configuration and cancellation
|
|
26
|
+
|
|
27
|
+
Repository acceptance: all 13 follow-up orders closed; 6,885 tests passed, 4 skipped across 641 suites; clean rebuild passed for all 11 workspace packages. Source head 7005aa41 and scoped repair commits are recorded in the program ledger. Publication and subsequent live Telegram acceptance remain with the user.
|
|
28
|
+
|
|
29
|
+
## September 4 field-review follow-up: WO-25 through WO-29
|
|
30
|
+
|
|
31
|
+
The active Telegram field review found new deterministic defects after the
|
|
32
|
+
earlier acceptance runs. Prior checked sections remain historical evidence;
|
|
33
|
+
they do not close these follow-up repairs. Current authority is
|
|
34
|
+
[TELEGRAM-LONGHAUL-2026-09-04.md](TELEGRAM-LONGHAUL-2026-09-04.md).
|
|
35
|
+
|
|
36
|
+
- [x] WO-25 memory compiler eligibility and outcome audit.
|
|
37
|
+
- [x] WO-26 admin Telegram media alias resolution.
|
|
38
|
+
- [x] WO-27 task verification authority separated from mutation observations.
|
|
39
|
+
- [x] WO-28 obsolete task evidence retirement after explicit replacement.
|
|
40
|
+
- [x] WO-29 retired failure diagnostics and active projections.
|
|
41
|
+
- [x] Aggregate tests, clean build, independent review, and scoped Git delivery.
|
|
42
|
+
- Publication and live validation are reserved to the user after handoff.
|
|
6
43
|
|
|
7
44
|
## Phase 0: durable program boundary
|
|
8
45
|
|
package/docs/work-orders/runtime-health-remediation/WO-25-memory-compiler-eligibility-and-audit.md
ADDED
|
@@ -0,0 +1,81 @@
|
|
|
1
|
+
# WO-25: Memory compiler eligibility and outcome audit
|
|
2
|
+
|
|
3
|
+
**Status:** complete; implemented, verified, and pushed to origin/main
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Incident date:** 2026-09-04 PDT / 2026-09-05 UTC
|
|
6
|
+
**Program:** [September 4 Telegram long-haul repairs](TELEGRAM-LONGHAUL-2026-09-04.md)
|
|
7
|
+
|
|
8
|
+
## Observed failure
|
|
9
|
+
|
|
10
|
+
In run telegram-64ac9937dc7647a8-1788572475259-1, 18 compiler operations consumed 5,371.58 seconds (89.53 minutes; 73.2% of 122.27 minutes). All 18 compiler prompts reported compactionEligible=false and only 20–30% occupancy of 262,144 tokens. A nineteenth was pending at the observation cutoff. These are completed operation durations, not inferred GPU compute or proven timeouts.
|
|
11
|
+
|
|
12
|
+
## Evidence
|
|
13
|
+
|
|
14
|
+
Paths are relative to /home/roko/Documents/Projects/Adjacent. Live files are
|
|
15
|
+
observational references; sanitized deterministic fixtures must reproduce the
|
|
16
|
+
failure without copying credentials or private conversation content.
|
|
17
|
+
|
|
18
|
+
- `telegram_test/.omnius/interruptions/telegram-64ac9937dc7647a8.json`
|
|
19
|
+
- `telegram_test/.omnius/context-window-dumps/2026-09-05T03-35-52-688Z-internal-3abab33a54.json`
|
|
20
|
+
|
|
21
|
+
## Root cause
|
|
22
|
+
|
|
23
|
+
_applyUnifiedMemoryCompilation bypasses context eligibility above 48,000 raw tool characters, synchronously requests a plan, and validateMemoryCompilationPlan rejects any such plan below the headroom threshold. Canonical projection subsequently changes the request fingerprint and drops the compiler audit.
|
|
24
|
+
|
|
25
|
+
## Code locations
|
|
26
|
+
|
|
27
|
+
- `packages/orchestrator/src/agenticRunner.ts — _applyUnifiedMemoryCompilation, _recordContextWindowDump`
|
|
28
|
+
- `packages/orchestrator/src/memory-compiler.ts — validation and analysis`
|
|
29
|
+
- `packages/orchestrator/src/contextWindowDump.ts — compilation audit contract`
|
|
30
|
+
- `packages/orchestrator/tests/unified-memory-compilation.test.ts`
|
|
31
|
+
- `packages/orchestrator/tests/memory-compiler.test.ts`
|
|
32
|
+
- `packages/orchestrator/tests/contextWindowDump.test.ts`
|
|
33
|
+
|
|
34
|
+
## Repair design
|
|
35
|
+
|
|
36
|
+
- Use one eligibility contract before inference and validation; raw history volume alone must not invoke a plan that cannot be accepted.
|
|
37
|
+
- Keep useful high-pressure compaction and graph/authority validation intact. Separate deterministic stale-evidence projection (WO-28) from semantic capacity compaction.
|
|
38
|
+
- Record compiler outcome against its actual input fingerprint and explicitly relate it to the final projected request, rather than silently dropping the audit.
|
|
39
|
+
- Preserve invalid-output/exception/hold outcomes as bounded body-free diagnostics without assuming every long request timed out.
|
|
40
|
+
|
|
41
|
+
## Acceptance checklist
|
|
42
|
+
|
|
43
|
+
- [x] A large raw-tool request with ample headroom makes zero compiler backend calls across changing turns.
|
|
44
|
+
- [x] An eligible high-pressure request still applies a valid plan while retaining authority and active evidence.
|
|
45
|
+
- [x] Invalid plans preserve the request and emit an inspectable outcome.
|
|
46
|
+
- [x] Compiler outcome remains inspectable when downstream projection changes request fingerprint.
|
|
47
|
+
- [x] Focused regression tests, affected typechecks, and integrated workspace build pass.
|
|
48
|
+
|
|
49
|
+
## Implementation and verification log
|
|
50
|
+
|
|
51
|
+
- Regression-first execution reproduced all three failures: four changing
|
|
52
|
+
low-pressure turns made four compiler calls, downstream projection lost the
|
|
53
|
+
outcome audit, and backend errors had no distinguishable cached outcome.
|
|
54
|
+
- The runner now uses the validator's exact compaction eligibility before
|
|
55
|
+
candidate persistence or backend invocation. The former raw-character
|
|
56
|
+
environment override no longer bypasses capacity eligibility.
|
|
57
|
+
- V2 analysis retains typed `plan`, `invalid_output`, and `backend_error`
|
|
58
|
+
outcomes without copying exception messages. Cached failures retain their
|
|
59
|
+
outcome and cache-hit status.
|
|
60
|
+
- Stage receipts retain pre/post compiler fingerprints plus the final wire
|
|
61
|
+
fingerprint and whether subsequent projection changed the payload. This
|
|
62
|
+
preserves observability without claiming cross-payload validation.
|
|
63
|
+
- `OMNIUS_SQLITE_TESTS=1 pnpm --filter @omnius/orchestrator exec vitest run
|
|
64
|
+
tests/unified-memory-compilation.test.ts tests/memory-compiler.test.ts
|
|
65
|
+
tests/contextWindowDump.test.ts tests/unified-memory-compilation-guard.test.ts
|
|
66
|
+
tests/agenticRunner-canonical-production.test.ts`: **46 tests passed**.
|
|
67
|
+
- `pnpm --filter @omnius/orchestrator typecheck`: **passed**.
|
|
68
|
+
- Source/test whitespace checks passed. Integrated build and independent review
|
|
69
|
+
will be recorded in the program verification ledger.
|
|
70
|
+
|
|
71
|
+
## Delivery
|
|
72
|
+
|
|
73
|
+
- Commits: `cb316de6` on `origin/main`.
|
|
74
|
+
- Clean `pnpm -r build` passed for all 11 workspace packages/apps. Final aggregate suite evidence is in the linked program ledger.
|
|
75
|
+
|
|
76
|
+
## Publication boundary
|
|
77
|
+
|
|
78
|
+
Repository repair and deterministic verification are this order's completion
|
|
79
|
+
boundary. The user will publish the package and then perform live validation.
|
|
80
|
+
No package publication, installed-package replacement, runtime-state repair,
|
|
81
|
+
or service restart is authorized by this order.
|
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
# WO-26: Admin Telegram media alias resolution
|
|
2
|
+
|
|
3
|
+
**Status:** complete; implemented, verified, and pushed to origin/main
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Incident date:** 2026-09-04 PDT / 2026-09-05 UTC
|
|
6
|
+
**Program:** [September 4 Telegram long-haul repairs](TELEGRAM-LONGHAUL-2026-09-04.md)
|
|
7
|
+
|
|
8
|
+
## Observed failure
|
|
9
|
+
|
|
10
|
+
The active run called transcribe_file(path='message_id:2747') at 19:03 PDT, received a literal filesystem-not-found error, and reached successful real-file transcription only at 19:58. Earlier aliases 'latest' and message_id:2744 failed similarly. The recovery delay overlaps WO-25 compiler time.
|
|
11
|
+
|
|
12
|
+
## Evidence
|
|
13
|
+
|
|
14
|
+
Paths are relative to /home/roko/Documents/Projects/Adjacent. Live files are
|
|
15
|
+
observational references; sanitized deterministic fixtures must reproduce the
|
|
16
|
+
failure without copying credentials or private conversation content.
|
|
17
|
+
|
|
18
|
+
- `telegram_test/.omnius/completion-ledgers/telegram-64ac9937dc7647a8-1788572475259-1.json`
|
|
19
|
+
|
|
20
|
+
## Root cause
|
|
21
|
+
|
|
22
|
+
Telegram media descriptions advertise aliases to all callers, but applyTelegramScopedMemoryTools installs alias adapters only outside telegram-admin-dm. Raw transcription resolves the alias against the workspace as a filename.
|
|
23
|
+
|
|
24
|
+
## Code locations
|
|
25
|
+
|
|
26
|
+
- `packages/cli/src/tui/telegram-bridge.ts — media descriptions, tool assembly, scoped media resolver`
|
|
27
|
+
- `packages/execution/src/tools/transcribe-tool.ts — filesystem path boundary (preserve generic tool semantics)`
|
|
28
|
+
- `packages/cli/tests/telegram-media-cache.test.ts`
|
|
29
|
+
- `packages/cli/tests/telegram-public-media.test.ts`
|
|
30
|
+
- `packages/cli/tests/telegram-media-evidence.test.ts`
|
|
31
|
+
- `packages/cli/tests/telegram-admin-media-aliases.test.ts`
|
|
32
|
+
|
|
33
|
+
## Repair design
|
|
34
|
+
|
|
35
|
+
- Separate Telegram reference resolution from public-chat filesystem/memory restrictions.
|
|
36
|
+
- Resolve message_id, latest, and reply aliases in admin DMs using the current chat's cached media.
|
|
37
|
+
- Preserve legitimate explicit admin filesystem paths and retain group/public isolation.
|
|
38
|
+
- Apply the adapter consistently to advertised media tools affected by the same missing boundary; preserve media evidence attribution.
|
|
39
|
+
|
|
40
|
+
## Acceptance checklist
|
|
41
|
+
|
|
42
|
+
- [x] Admin-DM message_id, latest, and reply references reach the selected cached file, never a literal alias filename.
|
|
43
|
+
- [x] Explicit permitted admin paths retain existing behavior.
|
|
44
|
+
- [x] Missing and cross-chat references fail precisely without exposing unrelated media or initiating broad filesystem search.
|
|
45
|
+
- [x] Public/group path restrictions and media-evidence attribution remain intact.
|
|
46
|
+
- [x] Focused media/transport tests and CLI typecheck pass.
|
|
47
|
+
|
|
48
|
+
## Implementation and verification log
|
|
49
|
+
|
|
50
|
+
- Added an admin-only alias adapter for image reading, OCR, vision, PDF extraction,
|
|
51
|
+
transcription, video understanding, and audio analysis. It reuses the existing
|
|
52
|
+
chat/kind selection and bounded reply-chain resolver independently of public
|
|
53
|
+
memory and filesystem restrictions.
|
|
54
|
+
- Added `video_understand`, `audio_analyze`, and `telegram_media_recent` to the
|
|
55
|
+
admin toolset because incoming-media instructions already advertise those
|
|
56
|
+
capabilities. Explicit tool-policy filtering remains in force.
|
|
57
|
+
- Exact missing or foreign-chat message aliases fail before raw execution or
|
|
58
|
+
creative-workspace path fallback. Admin absolute/relative files remain raw
|
|
59
|
+
filesystem inputs; an existing local filename takes precedence over an equal
|
|
60
|
+
cached basename. Public/group adapters retain their scope and restrictions.
|
|
61
|
+
- Preserved admin OCR batch/output arguments, URL-video operations, and explicit
|
|
62
|
+
microphone modes. Successful alias-based transcription and audio analysis use
|
|
63
|
+
the existing media-evidence attribution callbacks.
|
|
64
|
+
- Independent review found that omitted OCR batch input could select a cached
|
|
65
|
+
image instead of preserving the raw required-directory error. Batch mode now
|
|
66
|
+
bypasses all alias/default-media selection. A regression test compares the
|
|
67
|
+
original missing-input error with the adapted result while cached images exist.
|
|
68
|
+
- Regression evidence: the assembled-toolset test first executes the raw
|
|
69
|
+
transcription path with a nonexistent `message_id:10` filename and confirms
|
|
70
|
+
the former filesystem-not-found result. It then verifies that the repaired
|
|
71
|
+
real admin tool assembly sends the cached filename to a mocked executor.
|
|
72
|
+
The missing-file reproduction returns before any transcription/model request.
|
|
73
|
+
- Final verification on 2026-09-04 PDT: **71 tests passed** across
|
|
74
|
+
`telegram-admin-media-aliases.test.ts` (33), `telegram-media-evidence.test.ts`
|
|
75
|
+
(9), `telegram-media-cache.test.ts` (7), `telegram-public-media.test.ts` (11),
|
|
76
|
+
and `telegram-transport-contract.test.ts` (11), with the CLI hermetic network
|
|
77
|
+
boundary enabled. This final run includes the raw alias-boundary and missing
|
|
78
|
+
OCR batch-input regressions. `pnpm --dir packages/cli typecheck` and
|
|
79
|
+
`git diff --check` passed after the batch-input follow-up.
|
|
80
|
+
- No live inference, publication, installed-package replacement, runtime-state
|
|
81
|
+
change, or service restart was performed. Parent task owns the delivery commit.
|
|
82
|
+
|
|
83
|
+
## Delivery
|
|
84
|
+
|
|
85
|
+
- Commits: `bc494850` on `origin/main`.
|
|
86
|
+
- Clean `pnpm -r build` passed for all 11 workspace packages/apps. Final aggregate suite evidence is in the linked program ledger.
|
|
87
|
+
|
|
88
|
+
## Publication boundary
|
|
89
|
+
|
|
90
|
+
Repository repair and deterministic verification are this order's completion
|
|
91
|
+
boundary. The user will publish the package and then perform live validation.
|
|
92
|
+
No package publication, installed-package replacement, runtime-state repair,
|
|
93
|
+
or service restart is authorized by this order.
|
|
@@ -0,0 +1,140 @@
|
|
|
1
|
+
# WO-27: Separate mutation observations from task verification
|
|
2
|
+
|
|
3
|
+
**Status:** complete; implemented, verified, and pushed to origin/main
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Incident date:** 2026-09-04 PDT / 2026-09-05 UTC
|
|
6
|
+
**Program:** [September 4 Telegram long-haul repairs](TELEGRAM-LONGHAUL-2026-09-04.md)
|
|
7
|
+
|
|
8
|
+
## Observed failure
|
|
9
|
+
|
|
10
|
+
At 01:54 PDT September 4, scaffold run telegram-64ac9937dc7647a8-1788511366966-1 completed with npm install --no-audit --no-fund 2>&1 | tail -20 as its sole shell command. Its readiness had zero obligations, claims, and verifier receipts. No test/build was executed.
|
|
11
|
+
|
|
12
|
+
## Evidence
|
|
13
|
+
|
|
14
|
+
Paths are relative to /home/roko/Documents/Projects/Adjacent. Live files are
|
|
15
|
+
observational references; sanitized deterministic fixtures must reproduce the
|
|
16
|
+
failure without copying credentials or private conversation content.
|
|
17
|
+
|
|
18
|
+
- `telegram_test/.omnius/completion-finalizations/telegram-64ac9937dc7647a8-1788511366966-1.json`
|
|
19
|
+
- `telegram_test/.omnius/terminal-trajectories/ (matching run)`
|
|
20
|
+
|
|
21
|
+
## Root cause
|
|
22
|
+
|
|
23
|
+
postActionVerifier labels a filesystem mtime observation outcomeClass='verified'. The no-todo runner path accepts that as direct task validation, and completionAutoFinalize can then report that changed files were validated. Observation of a successful mutation is weaker evidence than verification of requested behavior.
|
|
24
|
+
|
|
25
|
+
## Code locations
|
|
26
|
+
|
|
27
|
+
- `packages/orchestrator/src/postActionVerifier.ts`
|
|
28
|
+
- `packages/orchestrator/src/taskValidationEvidence.ts`
|
|
29
|
+
- `packages/orchestrator/src/verificationInvocation.ts`
|
|
30
|
+
- `packages/orchestrator/src/verificationCommand.ts`
|
|
31
|
+
- `packages/orchestrator/src/agenticRunner.ts — no-todo validation recording and completion readiness`
|
|
32
|
+
- `packages/orchestrator/src/completionAutoFinalize.ts`
|
|
33
|
+
- `packages/orchestrator/tests/completionAutoFinalize.test.ts`
|
|
34
|
+
- `packages/orchestrator/tests/agenticRunner-completionReadiness.test.ts`
|
|
35
|
+
- `packages/orchestrator/tests/postActionVerifier.test.ts`
|
|
36
|
+
|
|
37
|
+
## Repair design
|
|
38
|
+
|
|
39
|
+
- Represent mutation observation and task verification distinctly at the producer/consumer boundary.
|
|
40
|
+
- Prevent install/create/edit success and recent mtimes from minting task-validation authority.
|
|
41
|
+
- Retain explicit declared verifiers and real successful no-todo test/build verification, scoped to the actual execution and current mutation revision.
|
|
42
|
+
- Keep read-only/informational completion possible without inventing test obligations.
|
|
43
|
+
- Require fresh verification after a subsequent mutation and preserve truthful terminal summaries.
|
|
44
|
+
|
|
45
|
+
## Acceptance checklist
|
|
46
|
+
|
|
47
|
+
- [x] Scaffold edits followed only by dependency installation do not auto-finalize as verified.
|
|
48
|
+
- [x] Mutation-only verifier observations cannot satisfy no-todo validation.
|
|
49
|
+
- [x] A declared verifier and a real successful no-todo test/build retain valid completion behavior.
|
|
50
|
+
- [x] Failed or stale verification cannot authorize completion; a later mutation invalidates earlier verification.
|
|
51
|
+
- [x] Read-only tasks do not require artificial build/test steps.
|
|
52
|
+
- [x] Focused completion/verifier tests and orchestrator typecheck pass.
|
|
53
|
+
|
|
54
|
+
## Implementation and verification log
|
|
55
|
+
|
|
56
|
+
- Post-action audits now explicitly carry `completionAuthority: "none"`.
|
|
57
|
+
Their existing `verified` outcome remains an audit finding for telemetry;
|
|
58
|
+
it cannot establish acceptance of the user's task.
|
|
59
|
+
- Added typed task-validation evidence sourced only from a successful declared
|
|
60
|
+
verifier or an assertive validation execution. The selector excludes runtime
|
|
61
|
+
replays, contradictory/mismatched receipts, masked failures, and test/build
|
|
62
|
+
keywords used only as package arguments or file paths.
|
|
63
|
+
- No-todo automatic completion requires the typed evidence to match the
|
|
64
|
+
recorded command and turn. Tool-call identity invalidates a preceding pass
|
|
65
|
+
after another tool mutates files in the same turn, while preserving a build
|
|
66
|
+
that writes its own verified artifacts. Later failed checks revoke prior
|
|
67
|
+
validation even when no intervening mutation occurred.
|
|
68
|
+
- Added a mocked runner fixture with no subprocess or network execution. The
|
|
69
|
+
initial integration run reproduced both installation commands minting
|
|
70
|
+
`_lastBuildSuccessTurn=0`, and a subsequent same-turn edit leaving that
|
|
71
|
+
validation intact. These three assertions failed before runner integration.
|
|
72
|
+
- The explicit custom verifier fixture also exposed an existing producer gap:
|
|
73
|
+
`test -f ...` matched the declared verifier but the generic classifier
|
|
74
|
+
dismissed it as a read-only command. Its successful execution must reach the
|
|
75
|
+
completion ledger as declared verification; read-only tasks themselves gain
|
|
76
|
+
no new test obligation.
|
|
77
|
+
- Runner integration now routes explicit declared verifiers into the assertion
|
|
78
|
+
evidence ledger, carries typed authority into both automatic finalization
|
|
79
|
+
boundaries, and invalidates stale authority after mutations or failed checks.
|
|
80
|
+
- Verification completed September 5, 2026 UTC, from `packages/orchestrator`:
|
|
81
|
+
- `pnpm exec vitest run tests/taskValidationEvidence.test.ts tests/completionAutoFinalize.test.ts tests/postActionVerifier.test.ts tests/agenticRunner-taskValidationAuthority.test.ts tests/agenticRunner-completionReadiness.test.ts tests/verificationCommand.test.ts tests/completion-provenance-visual-evidence.test.ts` — **92 passed across 7 files**.
|
|
82
|
+
- `pnpm exec vitest run tests/agenticRunner.test.ts -t 'synthesizes truth-based completion when final-turn validation passes'` — **1 passed; 174 unrelated cases skipped**.
|
|
83
|
+
- `pnpm exec tsc --noEmit` — passed.
|
|
84
|
+
- The runner regressions use mock backends and mock shell tools. No live model
|
|
85
|
+
request, package installation, service mutation, or installed-package change
|
|
86
|
+
was performed. Git delivery is recorded by the parent work-order owner.
|
|
87
|
+
|
|
88
|
+
### Independent-review followup after ce40bfca
|
|
89
|
+
|
|
90
|
+
- Review found that using the filesystem-audit intent parser also rejected real
|
|
91
|
+
commands with launcher options or executable paths, including `pnpm --filter
|
|
92
|
+
@omnius/cli build`, `npm --prefix app run build`, a local `.bin/vitest`, and
|
|
93
|
+
`npx --package typescript tsc --noEmit`. It also treated help/version and
|
|
94
|
+
collection modes as validation.
|
|
95
|
+
- Added a literal, quote-aware stage parser on the verification receipt
|
|
96
|
+
boundary and a separate invocation classifier. Launcher options consume
|
|
97
|
+
their values before action classification; executable basenames and inert
|
|
98
|
+
reporting suffixes are handled structurally. Unknown shell expansion stays
|
|
99
|
+
unclassified. Existing historical command matching remains unchanged.
|
|
100
|
+
- A matched declared command now has to verify the final state. For example,
|
|
101
|
+
`pnpm test && touch src/example.ts` cannot reuse the earlier test as proof of
|
|
102
|
+
the subsequent edit. `pnpm test && echo done` remains valid; redirected
|
|
103
|
+
reporting that writes a file does not count as an inert suffix.
|
|
104
|
+
- An untrusted later check or contradictory execution receipt revokes an
|
|
105
|
+
earlier pass even when the outer success flag is true. Positive protocol
|
|
106
|
+
text cannot override a nonzero exit receipt or restore that stale evidence.
|
|
107
|
+
- Final followup verification, from `packages/orchestrator`:
|
|
108
|
+
- `pnpm exec vitest run tests/taskValidationEvidence.test.ts tests/verificationInvocation.test.ts tests/completionAutoFinalize.test.ts tests/postActionVerifier.test.ts tests/agenticRunner-taskValidationAuthority.test.ts tests/agenticRunner-completionReadiness.test.ts tests/verificationCommand.test.ts tests/completion-provenance-visual-evidence.test.ts` — **168 passed across 8 files**, including **25 runner cases**.
|
|
109
|
+
- `pnpm exec tsc --noEmit` — passed.
|
|
110
|
+
- All review examples have deterministic regressions. Followup Git delivery is
|
|
111
|
+
pending the parent owner's scoped commit; no runtime publication is implied.
|
|
112
|
+
|
|
113
|
+
### Full-suite compatibility followup
|
|
114
|
+
|
|
115
|
+
- The stale-review reconciliation fixture depended on the removed generic
|
|
116
|
+
completion heuristic despite declaring no verifier or edited artifact. It
|
|
117
|
+
now supplies a completed todo with `pnpm test` and a coherent execution
|
|
118
|
+
receipt, asserts declared verification authority, and disables automatic
|
|
119
|
+
finalization only to reach the explicit terminal boundary under test.
|
|
120
|
+
- Independent review also reproduced rejected valid Vitest `--run` and
|
|
121
|
+
`--coverage` invocations. Vitest mode options now use their own arity table,
|
|
122
|
+
separate from package-launcher options. Collection still remains inspection
|
|
123
|
+
when mode flags precede `list`; a config file named `list` remains a value.
|
|
124
|
+
- Verification from `packages/orchestrator`:
|
|
125
|
+
- `pnpm exec vitest run tests/verificationInvocation.test.ts tests/run-observability.test.ts tests/agenticRunner-taskValidationAuthority.test.ts` — **93 passed across 3 files** (57 invocation, 10 observability, 26 runner cases).
|
|
126
|
+
- `pnpm exec tsc --noEmit` — passed.
|
|
127
|
+
- This followup uses mocked execution and adds no live inference or runtime
|
|
128
|
+
state changes. Parent work-order owner handles the scoped commit and push.
|
|
129
|
+
|
|
130
|
+
## Delivery
|
|
131
|
+
|
|
132
|
+
- Commits: `ce40bfca, cf966ea3, dc60e5bf` on `origin/main`.
|
|
133
|
+
- Clean `pnpm -r build` passed for all 11 workspace packages/apps. Final aggregate suite evidence is in the linked program ledger.
|
|
134
|
+
|
|
135
|
+
## Publication boundary
|
|
136
|
+
|
|
137
|
+
Repository repair and deterministic verification are this order's completion
|
|
138
|
+
boundary. The user will publish the package and then perform live validation.
|
|
139
|
+
No package publication, installed-package replacement, runtime-state repair,
|
|
140
|
+
or service restart is authorized by this order.
|