omnius 1.0.695 → 1.0.697
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/api/py-embed.js +4215 -2598
- package/dist/index.js +41242 -35198
- package/dist/library.js +8063 -6080
- package/dist/python-cuda-runtime.js +512 -0
- package/dist/update-worker.js +4250 -2633
- package/docs/DISCOVERY.json +802 -1
- package/docs/DISCOVERY.md +22 -1
- package/docs/guides/long-horizon-feature-workflow.md +42 -0
- package/docs/research/aiwg-long-horizon-feature-workflow.md +94 -0
- package/docs/work-orders/long-horizon-feature-delivery/WO-01-native-research-feature-workflow.md +122 -0
- package/docs/work-orders/runtime-health-remediation/TOOL-QUALITY-2026-09-05.md +84 -0
- package/docs/work-orders/runtime-health-remediation/TRACKER.md +48 -1
- package/docs/work-orders/runtime-health-remediation/WO-30-structured-tool-invocation.md +65 -0
- package/docs/work-orders/runtime-health-remediation/WO-31-complete-runtime-policy.md +62 -0
- package/docs/work-orders/runtime-health-remediation/WO-32-evidence-dependent-steering.md +55 -0
- package/docs/work-orders/runtime-health-remediation/WO-33-file-mutation-transactions.md +55 -0
- package/docs/work-orders/runtime-health-remediation/WO-34-shell-authority-and-results.md +46 -0
- package/docs/work-orders/runtime-health-remediation/WO-35-search-and-exploration-isolation.md +56 -0
- package/docs/work-orders/runtime-health-remediation/WO-36-media-evidence-integrity.md +44 -0
- package/docs/work-orders/runtime-health-remediation/WO-37-browser-and-process-lifecycle.md +86 -0
- package/docs/work-orders/runtime-health-remediation/WO-37-point-localization-cancellation.md +24 -0
- package/docs/work-orders/runtime-health-remediation/WO-38-tool-contract-preservation.md +83 -0
- package/docs/work-orders/runtime-health-remediation/WO-39-telegram-working-indicator.md +76 -0
- package/docs/work-orders/runtime-health-remediation/WO-40-web-content-and-crawl-receipts.md +44 -0
- package/docs/work-orders/runtime-health-remediation/WO-41-completion-evidence-consistency.md +45 -0
- package/docs/work-orders/runtime-health-remediation/WO-42-media-execution-and-configuration.md +79 -0
- package/docs/work-orders/runtime-health-remediation/WO-43-telegram-router-progress-boundary.md +71 -0
- package/docs/work-orders/runtime-health-remediation/WO-44-evidence-backed-terminal-results.md +56 -0
- package/docs/work-orders/runtime-health-remediation/WO-45-native-ollama-tool-contract.md +33 -0
- package/npm-shrinkwrap.json +5 -5
- package/package.json +1 -1
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# WO-39: Keep Telegram's working indicator alive through active work
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Date:** 2026-09-05 PDT
|
|
6
|
+
|
|
7
|
+
## Requirement and root cause
|
|
8
|
+
|
|
9
|
+
The user requires a visible working indicator throughout background work in
|
|
10
|
+
Telegram DMs and public/private groups. All three private reply paths cleared
|
|
11
|
+
their typing interval when the first native draft was acknowledged. A following
|
|
12
|
+
silent tool call or unchanged draft had no refresh, so Telegram's short-lived
|
|
13
|
+
typing action expired while work continued. Group paths already used recurring
|
|
14
|
+
typing, but shutdown could not directly reach a quick-chat timer once that timer
|
|
15
|
+
had transferred out of the staging map.
|
|
16
|
+
|
|
17
|
+
An initial transient typing failure also returned permission to continue
|
|
18
|
+
without creating a refresh timer, leaving admitted media/context preparation
|
|
19
|
+
without its working heartbeat.
|
|
20
|
+
|
|
21
|
+
## Implementation
|
|
22
|
+
|
|
23
|
+
- Keep one run-owned 3-second heartbeat through media preparation, model/tool
|
|
24
|
+
work, native previews, and final delivery. Native draft success never retires
|
|
25
|
+
that heartbeat; terminal reply cleanup does.
|
|
26
|
+
- Track active timer cleanups after staging transfers, attach the admitted
|
|
27
|
+
work's cancellation signal, and clear every active lease on bridge shutdown.
|
|
28
|
+
Cleanup removes both the timer and its abort listener.
|
|
29
|
+
- Reject late initial acknowledgements from cancelled/replaced work; continue
|
|
30
|
+
refreshing after transient failures. Preserve permanent write refusals and
|
|
31
|
+
reply authorization before visible activity.
|
|
32
|
+
- Coalesce refresh ticks while one transport call remains in flight and reuse
|
|
33
|
+
an active runner's lease when steering arrives.
|
|
34
|
+
- Space native drafts at 1.2 seconds so draft updates and 3-second heartbeats
|
|
35
|
+
fit together within the existing shared 40-per-30-second typing budget.
|
|
36
|
+
- Preserve numeric DM/group/topic routing and guest-query exclusions.
|
|
37
|
+
|
|
38
|
+
Timers express an admitted live process's ownership; they are not persisted or
|
|
39
|
+
revived for stale runs after restart. Visible refresh still depends on Telegram
|
|
40
|
+
accepting the API request and a responsive process/transport.
|
|
41
|
+
|
|
42
|
+
## Verification
|
|
43
|
+
|
|
44
|
+
`packages/cli/tests/telegram-working-indicator.test.ts` exercises the real task,
|
|
45
|
+
admin-chat, and quick-chat reply handlers with mocked inference/tool work and
|
|
46
|
+
mocked Telegram API responses. It verifies successful first drafts followed by
|
|
47
|
+
18 seconds of silence, delayed final delivery, public/private groups and topics,
|
|
48
|
+
media preparation, cancellation/shutdown before completion settles, late
|
|
49
|
+
acknowledgements, transient failures, coalescing, steering, and shared budget.
|
|
50
|
+
|
|
51
|
+
- **220 tests passed** across eight suites: working-indicator, transport-contract,
|
|
52
|
+
attention-evidence, api-governor, write-capability, bridge-lifecycle,
|
|
53
|
+
lifecycle-leak, and bot-api-10.
|
|
54
|
+
- After adding explicit permanent-refusal ownership cleanup and generation
|
|
55
|
+
failure coverage, **76 tests passed** across six affected transport/lifecycle
|
|
56
|
+
suites, including all **17 working-indicator tests**.
|
|
57
|
+
- CLI `tsc --noEmit` and scoped `git diff --check` passed. All tests use the
|
|
58
|
+
CLI hermetic network boundary and synthetic temporary workspaces.
|
|
59
|
+
|
|
60
|
+
## Delivery boundary
|
|
61
|
+
|
|
62
|
+
Scoped source/test/document commit only. No real Telegram requests, inference,
|
|
63
|
+
publication, service restart, or mutation of the running Telegram workspace.
|
|
64
|
+
|
|
65
|
+
Final ownership review also routed queue-retirement cleanup through the shared typing lease cleanup, releasing its abort listener and ownership-map entry as well as the interval. All 56 tests across working-indicator, write-capability, bridge-lifecycle, and lifecycle-leak passed; /tmp/omnius-wo39-cleanup.log. The documented cadence now matches the implemented three seconds.
|
|
66
|
+
|
|
67
|
+
## Repository closure — September 5
|
|
68
|
+
|
|
69
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
70
|
+
|
|
71
|
+
Scoped repair commits delivered to origin/main: 53216239, c17ae247. Publication and subsequent live acceptance remain with the user.
|
|
72
|
+
|
|
73
|
+
Source and integration locations:
|
|
74
|
+
|
|
75
|
+
- [packages/cli/src/tui/telegram-bridge.ts](../../../packages/cli/src/tui/telegram-bridge.ts)
|
|
76
|
+
- [packages/cli/tests/telegram-working-indicator.test.ts](../../../packages/cli/tests/telegram-working-indicator.test.ts)
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# WO-40: Preserve retrieved content and validate crawler responses
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Date:** 2026-09-05
|
|
6
|
+
|
|
7
|
+
## Reproduced failures
|
|
8
|
+
|
|
9
|
+
- `web_fetch` strips HTML from every MIME type. A plain-text comparison becomes
|
|
10
|
+
different code, and JSON containing `<Widget />` loses that value. Short
|
|
11
|
+
plain text also produces misleading advice that JavaScript rendering is needed.
|
|
12
|
+
- `web_crawl` reports success for exit-zero workers emitting empty or malformed
|
|
13
|
+
output instead of the promised result JSON.
|
|
14
|
+
|
|
15
|
+
## Repair and acceptance
|
|
16
|
+
|
|
17
|
+
- Strip markup only from declared HTML/XHTML; preserve other textual bodies,
|
|
18
|
+
whitespace, and characters in both fresh responses and cached retrievals.
|
|
19
|
+
- Apply the requested output limit before any short-page hint.
|
|
20
|
+
- Require a successful, structurally valid crawler payload and retrieved pages;
|
|
21
|
+
malformed worker output is an execution diagnostic, not source evidence.
|
|
22
|
+
- Bound crawler numeric options and preserve existing network egress policy.
|
|
23
|
+
- Verify with mocked fetch/process boundaries; no live requests or services.
|
|
24
|
+
|
|
25
|
+
## Verification
|
|
26
|
+
|
|
27
|
+
- `pnpm exec vitest run tests/web-fetch.test.ts tests/web-crawl-contract.test.ts tests/web-download.test.ts tests/web-search.test.ts`
|
|
28
|
+
from `packages/execution`: **45 passed**.
|
|
29
|
+
- `pnpm exec tsc --noEmit` from `packages/execution`: passed.
|
|
30
|
+
- Fetch regressions preserve comparisons, JSX, whitespace, and JSON on fresh
|
|
31
|
+
and cached reads. Crawl regressions reject empty, malformed, zero-page and
|
|
32
|
+
contradictory worker outcomes while retaining valid attribution.
|
|
33
|
+
- Parent owns integration and push; publication is excluded.
|
|
34
|
+
|
|
35
|
+
## Repository closure — September 5
|
|
36
|
+
|
|
37
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
38
|
+
|
|
39
|
+
Scoped repair commits delivered to origin/main: 9f96cfcb. Publication and subsequent live acceptance remain with the user.
|
|
40
|
+
|
|
41
|
+
Source and integration locations:
|
|
42
|
+
|
|
43
|
+
- [packages/execution/src/tools/web-fetch.ts](../../../packages/execution/src/tools/web-fetch.ts)
|
|
44
|
+
- [packages/execution/src/tools/web-crawl.ts](../../../packages/execution/src/tools/web-crawl.ts)
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# WO-41: Expose verification coverage and preserve typed completion authority
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Program:** [September 5 follow-up](TOOL-QUALITY-2026-09-05.md)
|
|
6
|
+
|
|
7
|
+
## Observed failure and architecture boundary
|
|
8
|
+
|
|
9
|
+
The monitored 1.0.695 run ended at 00:58:09 PDT on September 5. Its terminal record has status=completed and completionDisposition=ready, while ledgerStatus=incomplete_verification and the ledger contains verification_missing after two code mutations. Several check commands ended with a reporting echo; their final shell status alone cannot establish the check's success.
|
|
10
|
+
|
|
11
|
+
Completion obligations and observed verification coverage are separate dimensions. A generic mutation does not create a mandatory verifier requirement. Explicit expected-outcome, todo, workboard, or configured verifier requirements govern readiness. Prose claims and critic assertions do not manufacture completion obligations. The repair must make coverage visible without restoring a generic completion veto.
|
|
12
|
+
|
|
13
|
+
## Code repair
|
|
14
|
+
|
|
15
|
+
- packages/orchestrator/src/completionLedger.ts: derive stable observed verification facts from execution evidence, separately from the authority projection. Preserve negative typed receipts even when a wrapper reports success; only later matching successful verification discharges the failure.
|
|
16
|
+
- packages/orchestrator/src/completionFinalization.ts: persist verificationAudit status and gaps independently of terminal disposition; reject a completed record when supplied readiness explicitly denies successful completion.
|
|
17
|
+
- packages/orchestrator/src/agenticRunner.ts: expose audit coverage in terminal receipt text and preserve the existing explicit-obligation authority projection. Reconcile expected-outcome verification assessment with typed readiness facts.
|
|
18
|
+
- packages/orchestrator/tests/completion-observed-authority.test.ts: exercise real runner admission and finalization using mocked mutation/check tools, including masked shell success, explicit verifier obligations, and generic mutation without a mandatory verifier.
|
|
19
|
+
|
|
20
|
+
## Acceptance
|
|
21
|
+
|
|
22
|
+
- [x] Generic mutation without an explicit verifier may complete while retaining an unverified audit.
|
|
23
|
+
- [x] Missing or masked verification cannot satisfy an explicitly required verifier.
|
|
24
|
+
- [x] A fresh successful assertion-bearing check satisfies the corresponding typed expected effect.
|
|
25
|
+
- [x] Speculative prose claims and controller bookkeeping do not manufacture coverage gaps.
|
|
26
|
+
- [x] Later failed/timed-out/negative declared verification remains observable despite a positive wrapper result.
|
|
27
|
+
- [x] Focused and aggregate regression suites, independent review, clean build, scoped commits and delivery recorded.
|
|
28
|
+
|
|
29
|
+
## Verification and delivery
|
|
30
|
+
|
|
31
|
+
All reproductions use temporary state and mocked inference. Live workspace, installed package, and publication are outside this repair's execution boundary. Aggregate results and final delivery references are recorded in the program workorder.
|
|
32
|
+
|
|
33
|
+
Focused verification: 48 tests passed across observed completion authority, completion ledger, finalization, and readiness. The production cases cover explicit requirements versus generic audit coverage. Broader architecture regressions: 234 passed; two remaining context-admission failures belong to WO-31 and are tracked separately. Logs: /tmp/omnius-wo41-revised.log and /tmp/omnius-tool-quality-regressions.log.
|
|
34
|
+
|
|
35
|
+
## Repository closure — September 5
|
|
36
|
+
|
|
37
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
38
|
+
|
|
39
|
+
Scoped repair commits delivered to origin/main: 6ee251c3. Publication and subsequent live acceptance remain with the user.
|
|
40
|
+
|
|
41
|
+
Source and integration locations:
|
|
42
|
+
|
|
43
|
+
- [packages/orchestrator/src/completionLedger.ts](../../../packages/orchestrator/src/completionLedger.ts)
|
|
44
|
+
- [packages/orchestrator/src/completionFinalization.ts](../../../packages/orchestrator/src/completionFinalization.ts)
|
|
45
|
+
- [packages/orchestrator/src/agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts)
|
package/docs/work-orders/runtime-health-remediation/WO-42-media-execution-and-configuration.md
ADDED
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
# WO-42: Media execution and configuration boundaries
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
|
|
6
|
+
## Confirmed problems
|
|
7
|
+
|
|
8
|
+
- Microphone capture joins device/output arguments into synchronous shell commands. Shell metacharacters are executable, paths with spaces split, and a recording blocks the event loop for its requested duration. Recording is intentionally finite (up to 300 seconds); its deadline must allow that duration rather than impose an unrelated short limit.
|
|
9
|
+
- Transcription temporarily assigns the entire environment around an awaited shared Node bridge, and permanently assigns its Python/PATH selection. A hermetic overlapping-call reproduction observed request A using B's environment, request B using the parent's environment, and the parent left with A's environment.
|
|
10
|
+
- ComfyUI HTTP deadlines end when headers arrive. A hermetic headers-only response left the body pending beyond the configured deadline with its signal un-aborted. Video mux/thumbnail subprocesses have no deadline, output bound, or termination escalation.
|
|
11
|
+
- Web search advertises fully on-device execution while submitting queries to DuckDuckGo.
|
|
12
|
+
|
|
13
|
+
## Accepted repair
|
|
14
|
+
|
|
15
|
+
Use argument-vector asynchronous capture with duration-aware deadlines, cancellation, typed arguments and actual output validation. Keep interpreter/device configuration in a dedicated transcription child process, including bridge shutdown and bounded process lifetime. Keep Comfy response consumption inside its deadline with bounded bodies, and reuse the shared process lifecycle for video postprocessing. Describe search's network destination truthfully.
|
|
16
|
+
|
|
17
|
+
## Verification and scope
|
|
18
|
+
|
|
19
|
+
Use stubbed child processes, fake response bodies and temporary synthetic modules/files only. Do not capture live audio, perform inference, send network traffic, start services, or alter the installed runtime. Record focused test/typecheck results and scoped delivery below.
|
|
20
|
+
|
|
21
|
+
## Implementation and verification ledger
|
|
22
|
+
|
|
23
|
+
- Capture now supplies literal argv to asynchronous processes, including Pulse devices advertised by enumeration. Deadlines include the requested recording duration plus five seconds; cancellation reaches the shared process-group termination/escalation boundary. Numeric inputs are validated before dispatch. Fresh staging output is required before replacing the final recording; failed/empty captures preserve existing output and do not request ingestion. Level sampling removes temporary files and rejects empty/incomplete PCM. Enumeration distinguishes unavailable probes from a successful empty list.
|
|
24
|
+
- Transcribe-cli is resolved without importing its bridge into the parent. A dedicated Node worker receives a copied environment with its selected interpreter and device configuration, writes a typed result payload, shuts its bridge down, and is bounded to five minutes with process-group cleanup. Synthetic concurrent child modules prove both environment selections remain private and the parent's interpreter/PATH stay unchanged.
|
|
25
|
+
- Comfy response deadlines cover headers, success/error body consumption and cancellation. JSON bodies and downloaded video sizes are bounded; missing prompt IDs and empty video downloads fail explicitly. Mux and thumbnail processes have 120-second/30-second deadlines, bounded output and termination escalation; a late zero exit cannot erase a timeout.
|
|
26
|
+
- Search descriptions now disclose that queries are sent to DuckDuckGo over the network.
|
|
27
|
+
|
|
28
|
+
Before-repair hermetic reproductions captured the unquoted synchronous shell command, observed A/B/parent environment cross-contamination, and proved a Comfy headers-only response clears its deadline before body completion. No shell metacharacters were executed and no real capture, inference, network request or service launch occurred.
|
|
29
|
+
|
|
30
|
+
Focused verification: 81 tests passed across audio-capture-lifecycle, transcribe-cli-worker, comfy-response-lifecycle, media-worker-lifecycle, video-generate, transcribe-tool-artifacts, transcribe-python-runtime and web-search. Execution package `tsc --noEmit` and scoped `git diff --check` passed. New worker tests execute only temporary synthetic JavaScript modules, never Python or model code. The parent coordinates aggregate review and pushing the scoped commit.
|
|
31
|
+
|
|
32
|
+
## Nested audio generation Stop follow-up
|
|
33
|
+
|
|
34
|
+
Video's optional audio generation used `AudioGenerateTool`, which had no cancellation hook and a separate worker implementation that returned on timeout before its child had drained. Audio generation now owns an invocation controller and exposes `cancel()`. The signal reaches Python import/runtime compatibility probes, environment creation, package installation, prewarm, and generation. The shared process runner bounds output, terminates the owned process group, escalates resistant workers, and waits for closure. A timeout or output overflow stays a failure even when shutdown exits zero.
|
|
35
|
+
|
|
36
|
+
The same controller remains active through resource admission and lease release. Stop ends both generation/prewarm fallback ladders, releases an admission that completes after cancellation, and prevents another worker from starting. A second call on the same tool instance is rejected while this cleanup is pending. Existing deadline/download-stall fallback behavior is retained for ordinary backend failures; user cancellation ends the invocation.
|
|
37
|
+
|
|
38
|
+
Focused verification: **40 tests passed across six audio/Playwright suites**, including real synthetic Node workers that resist TERM (never Python/model workers), pre-aborted dispatch, zero-exit overflow, both canceled fallback ladders, delayed admission cleanup, browser isolation, and screenshot conversion boundaries. Execution `tsc --noEmit` and scoped `git diff --check` passed. Parent owns aggregate review/push; VideoGenerateTool's nested signal binding is delivered separately by its owner.
|
|
39
|
+
|
|
40
|
+
## Invocation cancellation follow-up
|
|
41
|
+
|
|
42
|
+
Independent review found that several bounded worker helpers accepted an abort signal while their tool entrypoints never supplied one. Stop therefore did not reach an isolated transcription process, a video worker, or Comfy polling.
|
|
43
|
+
|
|
44
|
+
- File and URL transcription now own an invocation controller, reject overlapping use of the same instance, propagate cancellation through interpreter selection, package commands, isolated bridge/CLI/managed workers, URL body consumption and nested transcription, and prevent fallback dispatch after cancellation. Temporary URL inputs use a UUID to keep cleanup scoped to their owner.
|
|
45
|
+
- Video cancellation reaches Diffusers setup/import/install/generation, ffmpeg mux/thumbnail work, Comfy startup/request/body/polling, and nested audio generation. Cancelled candidate attempts release their broker lease and do not start the next model. A completed video remains in the mutation receipt when later audio work is cancelled.
|
|
46
|
+
- The shared process runner refuses already-aborted requests before spawn. Newly created Comfy services reuse the owned-service startup primitive; startup cancellation and failure drain only their owned process group.
|
|
47
|
+
- Cancellation of a client request does **not** establish cancellation of a workflow already accepted by an external/shared Comfy server. No global interrupt or service-wide kill is issued. If submission may have reached that server, the cancellation result states that server work may continue. This preserves other clients' ownership.
|
|
48
|
+
|
|
49
|
+
Verification: **87 tests passed across 13 combined browser/media suites**; the final URL download and nested-worker regressions expanded transcription invocation coverage to **4 passing tests**. Execution typecheck and scoped diff checks passed. Tests use mocked transport/processes and temporary synthetic JavaScript modules; no Python/model execution, live services, network downloads or publication occurred.
|
|
50
|
+
|
|
51
|
+
## Managed Whisper synchronous preparation closure
|
|
52
|
+
|
|
53
|
+
The final review found that managed Whisper still called `ensureCudaPythonVenvSync` on the daemon thread before its readiness check. That unchanged CUDA helper can perform synchronous 60-second probes, 120-second venv creation and 20-minute package installations. An AbortSignal checked around that call cannot be delivered while the event loop is blocked.
|
|
54
|
+
|
|
55
|
+
Managed Whisper now invokes the exact existing preparation module in an owned asynchronous Node child. Its copied environment preserves CUDA selection and preparation policy; options travel through stdin, and the child/descendant process group is bounded and canceled through `runProcessBuffer`. The result receipt contains only the interpreter and protected-Torch pip constraint. The parent reconstructs the environment locally, validates the receipt, preserves bounded setup failure diagnostics, and checks Stop before any subsequent import, package install or inference. No CUDA policy or broker rule was rewritten.
|
|
56
|
+
|
|
57
|
+
Verification: **27 tests passed across five transcription suites**, including a synthetic preparation module blocked in synchronous work while the parent timer delivers Stop, worker drain, pre-aborted dispatch, a preparation deadline, configuration preservation, failure receipts and the real `TranscribeFileTool` cancellation boundary. Execution typecheck and scoped diff checks passed. Temporary bundled-helper fixtures independently verified default sibling-module resolution for both `execution/dist/transcribe-cuda-preparation.js` and published `dist/index.js`.
|
|
58
|
+
|
|
59
|
+
The parent added the standalone `python-cuda-runtime.js` publish entry and both artifact audits. Its actual build configuration produced a 20,494-byte temporary module exporting the preparation function without invoking it or emitting a sourcemap; all 16 publish-policy tests passed (`/tmp/omnius-cuda-worker-package.log`, `/tmp/omnius-cuda-worker-package-tests.log`). Packaging changes are committed separately by the parent. No live Python/CUDA preparation, model loading, network request, installed runtime change or modification of the dirty publish staging directory occurred during this verification.
|
|
60
|
+
|
|
61
|
+
## Repository closure — September 5
|
|
62
|
+
|
|
63
|
+
All confirmed findings and independent review follow-ups in this order are resolved. The [program verification ledger](TOOL-QUALITY-2026-09-05.md#verification-ledger) records the final full suites and clean build. Earlier pending integration notes above are historical and superseded by this closure.
|
|
64
|
+
|
|
65
|
+
Scoped repair commits delivered to origin/main: 7bc9106d, 52cb2623, b82f88a9, ed341587, 793c5302. Publication and subsequent live acceptance remain with the user.
|
|
66
|
+
|
|
67
|
+
Source and integration locations:
|
|
68
|
+
|
|
69
|
+
- [packages/execution/src/tools/audio-capture.ts](../../../packages/execution/src/tools/audio-capture.ts)
|
|
70
|
+
- [packages/execution/src/tools/audio-generate.ts](../../../packages/execution/src/tools/audio-generate.ts)
|
|
71
|
+
- [packages/execution/src/tools/transcribe-tool.ts](../../../packages/execution/src/tools/transcribe-tool.ts)
|
|
72
|
+
- [packages/execution/src/tools/video-generate.ts](../../../packages/execution/src/tools/video-generate.ts)
|
|
73
|
+
- [packages/execution/src/tools/web-search.ts](../../../packages/execution/src/tools/web-search.ts)
|
|
74
|
+
- [packages/execution/src/transcribe-cli-worker.ts](../../../packages/execution/src/transcribe-cli-worker.ts)
|
|
75
|
+
- [packages/execution/src/transcribe-cuda-preparation.ts](../../../packages/execution/src/transcribe-cuda-preparation.ts)
|
|
76
|
+
- [packages/execution/src/transcribe-python-runtime.ts](../../../packages/execution/src/transcribe-python-runtime.ts)
|
|
77
|
+
- [scripts/build-publish.mjs](../../../scripts/build-publish.mjs)
|
|
78
|
+
- [scripts/audit-publish-artifacts.mjs](../../../scripts/audit-publish-artifacts.mjs)
|
|
79
|
+
- [scripts/tarball-audit.mjs](../../../scripts/tarball-audit.mjs)
|
package/docs/work-orders/runtime-health-remediation/WO-43-telegram-router-progress-boundary.md
ADDED
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
# WO-43: Keep router failures out of Telegram progress and repair routing contracts
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; verified for user publication
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Reported:** September 5, 2026, after publication of 1.0.696
|
|
6
|
+
**Source baseline:** d8e6e2ed on origin/main
|
|
7
|
+
|
|
8
|
+
## User-visible failure
|
|
9
|
+
|
|
10
|
+
The initial Telegram card displays Working / Understanding followed by internal typed-router failure, retry and admin-admission diagnostics. The user reports seeing this first on every interaction. This is generated by the host rather than an assistant-authored explanation.
|
|
11
|
+
|
|
12
|
+
Read-only process inspection confirms the active Telegram workspace is using installed Omnius 1.0.696. No installed files or running services were changed.
|
|
13
|
+
|
|
14
|
+
## Root cause and code locations
|
|
15
|
+
|
|
16
|
+
The router's diagnostic reason is copied into subAgent.intakeComprehension when no explicit expected-outcome reason exists. The admin panel then renders that field as Understanding and publishes it during initial startup status handling.
|
|
17
|
+
|
|
18
|
+
- [telegram-bridge.ts](../../../packages/cli/src/tui/telegram-bridge.ts): processTelegramMessageWork, buildTelegramRouterUnavailableDecision, ensureTelegramAdminLivePanel, renderTelegramAdminWorkingSummary.
|
|
19
|
+
- [telegram-admin-live-panel.test.ts](../../../packages/cli/tests/telegram-admin-live-panel.test.ts) and [telegram-bot-api-10.test.ts](../../../packages/cli/tests/telegram-bot-api-10.test.ts): existing summary and router tests lack this production failure-to-progress boundary.
|
|
20
|
+
- [telegram-working-indicator.test.ts](../../../packages/cli/tests/telegram-working-indicator.test.ts): ongoing three-second typing ownership must remain intact.
|
|
21
|
+
|
|
22
|
+
An earlier persisted decision at 00:58 PDT on September 5 records router JSON admission rejected with logical_request_conflict, followed by a plain response and a strict retry rejected with direct_turn_conflicts_with_reply_target. Read-only source analysis reproduced the matching defects:
|
|
23
|
+
|
|
24
|
+
- Generated broker request identity hashes only sessionKey plus an instance-local inf-N counter. Restarting the bridge resets the counter, so the same session and inference kind can reuse an old broker identity for different request content. A per-bridge cryptographic nonce must separate lifetimes while each admitted logical request retains its identity across transport retries.
|
|
25
|
+
- The host evidence packet correctly marks private messages as directDeliveryToSelf. However, both the validator and routing prompt reject a direct turn when its reply edge targets another actor, without a private-delivery exception. An operator replying to their own earlier message in a DM is still speaking directly to the bot. The repair must honor that host-owned transport fact while preserving group reply-target checks and explicit self/evidence requirements.
|
|
26
|
+
|
|
27
|
+
The historical receipt establishes concrete matching failure mechanisms, not a claim that every reported failure had the same cause. The latest running process is 1.0.696; no live inference was sent to reproduce it.
|
|
28
|
+
|
|
29
|
+
## Repair plan and acceptance
|
|
30
|
+
|
|
31
|
+
- [x] Trace the reported string to its exact construction and delivery path.
|
|
32
|
+
- [x] Initialize request comprehension only from validated expected-outcome comprehension, never from routing reasons.
|
|
33
|
+
- [x] Keep internal reasons and detailed failure receipts available in diagnostics; ordinary progress remains useful and plain.
|
|
34
|
+
- [x] Cover unavailable and valid routers without outcome summaries, plus genuine outcome-summary display.
|
|
35
|
+
- [x] Reproduce and resolve confirmed request-correlation or reply-evidence contradictions without weakening typed reply authorization.
|
|
36
|
+
- [x] Preserve background typing, native draft lifecycle, cancellation and final delivery.
|
|
37
|
+
- [x] Complete independent review, affected tests, build and scoped commit/push to origin/main.
|
|
38
|
+
|
|
39
|
+
## Implementation and review
|
|
40
|
+
|
|
41
|
+
Routing reasons remain in TUI and durable social decision diagnostics. The initial Telegram response panel uses only comprehension from a validated, identity-bound expected-outcome contract. Otherwise its existing Working / Intake / Accepted presentation remains available. Native drafts and the independently owned typing heartbeat retain their lifecycle.
|
|
42
|
+
|
|
43
|
+
Generated inference identities include a cryptographic namespace for the bridge lifetime. Transport retries preserve the identity of the same request; a newly constructed bridge cannot reuse the previous bridge's inf-N identity in the same session.
|
|
44
|
+
|
|
45
|
+
Both initial and recovery prompts distinguish private transport delivery from group reply edges. The validator accepts private current-message evidence for direct delivery, while retaining self-role, addressed-actor and citation validation. Group replies aimed at someone else still require the appropriate authorization basis.
|
|
46
|
+
|
|
47
|
+
Independent review also found that failed normalization or rebinding could preserve a raw expected-outcome contract and stale matching decision IDs. This was found in synthetic boundary tests, not established as the live router failure. Invalid contracts are now cleared after safe effect augmentation/normalization, and failed trusted rebinding removes both the rejected contract and its stale decision ID. Presentation independently requires matching input and decision receipts. Existing implied visible-response effects are preserved.
|
|
48
|
+
|
|
49
|
+
## Runtime and publication boundary
|
|
50
|
+
|
|
51
|
+
Tests use temporary fixtures, mocked inference and mocked Telegram transport. No live inference, Telegram messages, GPU workloads, service restart or package publication is authorized by this repair. Existing publication staging and unrelated discovery changes are preserved. The user owns publication and subsequent live validation.
|
|
52
|
+
|
|
53
|
+
## Verification and delivery
|
|
54
|
+
|
|
55
|
+
- Baseline reproductions fail with the exact synthetic broker 409 identity conflict and with direct-turn conflicts for private replies to two different author identities. The initial recovery fixture had a separate mock setup error; it was corrected before final verification. No live request was involved.
|
|
56
|
+
- Presentation: 35 tests passed across the new seven-case production intake presentation suite, admin live panel and working indicator; four selected Bot API regressions also passed. The tests assert actual mocked first-message payloads, retained TUI/durable diagnostics, nine seconds of typing refresh, final response and cleanup.
|
|
57
|
+
- Routing: 29 tests passed across the ten new identity/private-reply cases, attention evidence and inference contracts. Coverage includes fresh bridge identity, stable queue retries, group rejection controls, strict recovery, malformed outcomes and stale outcome rebinding.
|
|
58
|
+
- Independent review resolved the stale receipt counterexample and confirmed group authorization, supplied logical IDs, unary request variants and typing/native draft ownership remain intact. Parent review confirmed safe array-shape handling preserves the existing implied visible-response effect before final normalization.
|
|
59
|
+
- CLI clean rebuild passed. The complete CLI suite passed all **2,531 tests across 268 suites**, with zero failures or skipped tests. This includes Telegram intake/typing/delivery and the daemon frontend transport regression suites.
|
|
60
|
+
|
|
61
|
+
Local reproduction and focused logs: /tmp/omnius-telegram-router-roots-before.log and /tmp/omnius-telegram-router-roots-after.log. Final logs: /tmp/omnius-wo43-cli-build.log and /tmp/omnius-wo43-cli-verified.log.
|
|
62
|
+
|
|
63
|
+
Scoped repair commit: **3f3c37f1**, delivered to **origin/main** with this workorder closure. The installed package and publication staging remain untouched; live acceptance follows the user's publication.
|
|
64
|
+
|
|
65
|
+
Regression source: [telegram-intake-presentation.test.ts](../../../packages/cli/tests/telegram-intake-presentation.test.ts) and [telegram-router-identity-and-private-replies.test.ts](../../../packages/cli/tests/telegram-router-identity-and-private-replies.test.ts).
|
|
66
|
+
|
|
67
|
+
## Field acceptance after user publication
|
|
68
|
+
|
|
69
|
+
Start a DM task without an explicit outcome summary: the initial card may show Working / Intake / Accepted, but no routing reason or failure prose. Reply to an earlier own message and confirm direct delivery is accepted. Typing continues throughout silent work, and genuine task summaries can replace initial intake. Restart the bridge and confirm its next inference is admitted without a reused logical-request conflict. Group replies to other actors retain their normal authorization checks.
|
|
70
|
+
|
|
71
|
+
These field checks remain pending user publication; repository tests use mocked transport.
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
# WO-44: Deliver an evidence-backed final result in Telegram
|
|
2
|
+
|
|
3
|
+
**Status:** complete in repository; publication and live acceptance remain with the user
|
|
4
|
+
**Priority:** P1
|
|
5
|
+
**Source baseline:** 44fa69cc
|
|
6
|
+
**Observed runtime:** 1.0.696, sibling telegram_test; read-only inspection
|
|
7
|
+
|
|
8
|
+
## Reported failure
|
|
9
|
+
|
|
10
|
+
After applying the operator's continuation of boutique-agent-services, Telegram ended with: “The run reached its completion boundary, but it did not produce a user-facing result. Open Evidence for the recorded outcome.”
|
|
11
|
+
|
|
12
|
+
Run telegram-64ac9937dc7647a8-1788624121778-1 finalized at 09:27:18 PDT on September 5 with status completed, disposition ready, and task epoch 1. Its host summary records four changed files (payment.ts, solana.ts, signature.ts and boutique.test.ts) and a successful post-mutation TypeScript check. The preceding tool output reports 21 passing tests; its piped shell receipt is classified as observation, so it must not be silently upgraded into typed verifier authority.
|
|
13
|
+
|
|
14
|
+
## Confirmed root causes
|
|
15
|
+
|
|
16
|
+
- Truth-based automatic completion emits assistant_text with source task_complete_summary and finishes without a separately authored user_reply. Telegram intentionally discards untyped runner summaries to avoid exposing bookkeeping. No structured result crosses that gap.
|
|
17
|
+
- A typed model_visible_text answer delivered without stream events can also be lost: the retention helper considers stream/accumulated content, and completion then excludes uncommitted assistant text.
|
|
18
|
+
- Auxiliary handoff grounding occurs after terminal commit and writes memory handoff state; its advisory outcome is not a user-facing terminal result.
|
|
19
|
+
- Finalization's selected command evidence excludes failed observations, while testsRun is attempted-command metadata. Neither can independently establish a complete, truthful check report.
|
|
20
|
+
|
|
21
|
+
## Code and contracts
|
|
22
|
+
|
|
23
|
+
- [completionAutoFinalize.ts](../../../packages/orchestrator/src/completionAutoFinalize.ts): implicit completion from direct file and validation evidence.
|
|
24
|
+
- [agenticRunner.ts](../../../packages/orchestrator/src/agenticRunner.ts): terminal commit, emitted events, AgenticResult and auxiliary handoff.
|
|
25
|
+
- [completionFinalization.ts](../../../packages/orchestrator/src/completionFinalization.ts): immutable terminal receipt and selected projections.
|
|
26
|
+
- [telegram-bridge.ts](../../../packages/cli/src/tui/telegram-bridge.ts): selectTelegramFinalResponse, runSubAgent, task_complete adapter and all final delivery paths.
|
|
27
|
+
|
|
28
|
+
## Accepted design and work plan
|
|
29
|
+
|
|
30
|
+
- [x] Reconstruct the reported terminal path from live receipts and source.
|
|
31
|
+
- [x] Build a bounded host-owned terminal report from current-epoch typed ledger evidence at commit, bound to run, epoch and terminal receipt.
|
|
32
|
+
- [x] Preserve actual file effects, check outcomes/freshness, failures and explicit gaps. Report truncation rather than implying complete coverage.
|
|
33
|
+
- [x] Deliver the typed report when no accepted user-facing answer exists; retain explicit user replies and genuine non-stream model answers.
|
|
34
|
+
- [x] Reject stale reports, unaccepted completion replies and arbitrary bookkeeping/old streams.
|
|
35
|
+
- [x] Replace opaque final fallbacks with a useful result or a clear account of missing evidence and the next required action.
|
|
36
|
+
- [x] Preserve typing, Stop, authenticated delivery, artifact receipts and inert-link validation.
|
|
37
|
+
- [x] Complete production-path mocked regressions, independent review, build and scoped commit/push.
|
|
38
|
+
|
|
39
|
+
Completion authority and verification audit remain separate. A generic mutation does not manufacture mandatory checks. The final result reports the evidence that exists, including uncertainty; it does not promote attempted commands or model claims to verified outcomes.
|
|
40
|
+
|
|
41
|
+
## Relationship to long-horizon work
|
|
42
|
+
|
|
43
|
+
The user additionally requested adaptation of AIWG's research-team mechanism for feature specification, creation, integration and testing across very long runs. Its source mechanisms and fit with existing Omnius workboard/recovery are recorded in [the research mapping](../../research/aiwg-long-horizon-feature-workflow.md), with native implementation tracked in [workflow WO-01](../long-horizon-feature-delivery/WO-01-native-research-feature-workflow.md). This terminal report provides the durable, evidence-bound handoff those longer workflows need; it does not by itself close that broader request.
|
|
44
|
+
|
|
45
|
+
## Verification and delivery
|
|
46
|
+
|
|
47
|
+
Implementation: 691532df, delivered to origin/main with this closure record.
|
|
48
|
+
|
|
49
|
+
- Full orchestrator regression: 218 suites passed; 2,650 tests passed and one intentionally skipped with OMNIUS_SQLITE_TESTS=1.
|
|
50
|
+
- Full CLI regression: 269 suites and 2,550 tests passed. An initial run found one source-regex fixture expecting the prior event guard; the expectation was updated for the stronger terminal-settled guard, its 51 focused tests passed, and the full CLI suite was rerun successfully.
|
|
51
|
+
- Orchestrator and CLI builds passed; git diff --check passed. Independent review covered actual mutation attribution, typed command role and freshness, failed or missing checks, held completion replies, epoch/receipt isolation, and bounded Telegram payloads.
|
|
52
|
+
- Production runner tests in accepted-terminal-reply.test.ts cover accepted replies, held attempts, terminal readiness rejection and runner reuse. terminal-task-report-runner.test.ts covers report commit and event/result identity. telegram-terminal-delivery.test.ts exercises actual mocked bridge delivery, explicit/model/report precedence, stale ownership, missing evidence and long-result truncation with full Evidence detail.
|
|
53
|
+
|
|
54
|
+
The report is generated from the committed ledger without another model request. Telegram receives concrete changed paths, recorded checks, limitations and next actions; untyped piped command output is never upgraded to a verification pass.
|
|
55
|
+
|
|
56
|
+
Tests use mocked inference/Telegram and temporary fixtures. No live inference, service restart, installed package replacement, external messaging or publication occurred. Existing publication staging and discovery changes remain untouched. The installed runtime remains unchanged until the user publishes.
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
# WO-45: Preserve tools across the native Ollama transport
|
|
2
|
+
|
|
3
|
+
**Status:** complete — implementation, mocked transport verification and scoped Git delivery
|
|
4
|
+
**Priority:** P1 for callers selecting native Ollama chat with tools
|
|
5
|
+
**Source baseline:** af6fb8eb
|
|
6
|
+
|
|
7
|
+
## Finding
|
|
8
|
+
|
|
9
|
+
Tracing the new feature workflow schema through the actual backend found that its definitions and references survive the ordinary OpenAI-compatible and Ollama-v1 transports. The separate native Ollama unary and streaming request builders in `packages/orchestrator/src/agenticRunner.ts` omit the request's tools entirely.
|
|
10
|
+
|
|
11
|
+
This is a transport contract defect for native-chat callers, not evidence that the normal v1 feature-workflow path loses its schema. A tool-capable request cannot work reliably if the adapter silently removes every tool before sending it.
|
|
12
|
+
|
|
13
|
+
## Repair and acceptance
|
|
14
|
+
|
|
15
|
+
- [x] Reproduce the native unary and streaming omissions with mocked HTTP responses.
|
|
16
|
+
- [x] Preserve the exact advertised tools, including nested definitions and references, in both native request bodies.
|
|
17
|
+
- [x] Preserve returned native tool calls and their arguments through the normalized runner response and stream lifecycle.
|
|
18
|
+
- [x] Retain direct-answer/tool thinking policy, cancellation and ordinary text-only behavior.
|
|
19
|
+
- [x] Run focused transport regressions, package checks and scoped Git delivery; record the implementation and verification here.
|
|
20
|
+
|
|
21
|
+
The actual host runner and HTTP encoders are in scope. No live model loading, inference, service restart, installed-package replacement or publication is authorized by this verification task.
|
|
22
|
+
|
|
23
|
+
## Verification and delivery — 2026-09-05
|
|
24
|
+
|
|
25
|
+
Commit **26d4cbed** is delivered to **origin/main**. The repair adds the original nonempty `request.tools` to both native request bodies in `agenticRunner.ts`; existing native response decoders already preserve typed tool calls and required no change.
|
|
26
|
+
|
|
27
|
+
`packages/orchestrator/tests/feature-workflow-transport.test.ts` contributes eight mocked transport cases. Exact tool schemas, including definitions and references, are compared at the actual HTTP boundary for ordinary compatible requests and native unary/streaming requests; native tool responses retain their names and arguments. The focused transport set passed **38 tests**. The final orchestrator suite passed **2,736 tests**, with **one skipped**, across **224 suites**; the clean workspace build and final workspace rebuild passed. Full cross-package acceptance is recorded in [workflow WO-01](../long-horizon-feature-delivery/WO-01-native-research-feature-workflow.md).
|
|
28
|
+
|
|
29
|
+
Publication and live model compatibility testing remain with the user.
|
|
30
|
+
|
|
31
|
+
## Related work
|
|
32
|
+
|
|
33
|
+
Discovered while verifying production tool delivery for [native feature workflow WO-01](../long-horizon-feature-delivery/WO-01-native-research-feature-workflow.md). The preceding Telegram final-result repair is [WO-44](WO-44-evidence-backed-terminal-results.md).
|
package/npm-shrinkwrap.json
CHANGED
|
@@ -1,12 +1,12 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "omnius",
|
|
3
|
-
"version": "1.0.
|
|
3
|
+
"version": "1.0.697",
|
|
4
4
|
"lockfileVersion": 3,
|
|
5
5
|
"requires": true,
|
|
6
6
|
"packages": {
|
|
7
7
|
"": {
|
|
8
8
|
"name": "omnius",
|
|
9
|
-
"version": "1.0.
|
|
9
|
+
"version": "1.0.697",
|
|
10
10
|
"bundleDependencies": [
|
|
11
11
|
"image-to-ascii"
|
|
12
12
|
],
|
|
@@ -5161,9 +5161,9 @@
|
|
|
5161
5161
|
}
|
|
5162
5162
|
},
|
|
5163
5163
|
"node_modules/jose": {
|
|
5164
|
-
"version": "6.2.
|
|
5165
|
-
"resolved": "https://registry.npmjs.org/jose/-/jose-6.2.
|
|
5166
|
-
"integrity": "sha512-
|
|
5164
|
+
"version": "6.2.12",
|
|
5165
|
+
"resolved": "https://registry.npmjs.org/jose/-/jose-6.2.12.tgz",
|
|
5166
|
+
"integrity": "sha512-9NiFmJEex0sy2Dk58j2UGBSHgUs2ypF9eZSu4L6vjOX3Dp96Sw1F3uL+H+D1sx02jZZdzUT0HgvCy59CuvXcWw==",
|
|
5167
5167
|
"license": "MIT",
|
|
5168
5168
|
"funding": {
|
|
5169
5169
|
"url": "https://github.com/sponsors/panva"
|
package/package.json
CHANGED