omnius 1.0.694 → 1.0.695

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,140 @@
1
+ # WO-27: Separate mutation observations from task verification
2
+
3
+ **Status:** complete; implemented, verified, and pushed to origin/main
4
+ **Priority:** P1
5
+ **Incident date:** 2026-09-04 PDT / 2026-09-05 UTC
6
+ **Program:** [September 4 Telegram long-haul repairs](TELEGRAM-LONGHAUL-2026-09-04.md)
7
+
8
+ ## Observed failure
9
+
10
+ At 01:54 PDT September 4, scaffold run telegram-64ac9937dc7647a8-1788511366966-1 completed with npm install --no-audit --no-fund 2>&1 | tail -20 as its sole shell command. Its readiness had zero obligations, claims, and verifier receipts. No test/build was executed.
11
+
12
+ ## Evidence
13
+
14
+ Paths are relative to /home/roko/Documents/Projects/Adjacent. Live files are
15
+ observational references; sanitized deterministic fixtures must reproduce the
16
+ failure without copying credentials or private conversation content.
17
+
18
+ - `telegram_test/.omnius/completion-finalizations/telegram-64ac9937dc7647a8-1788511366966-1.json`
19
+ - `telegram_test/.omnius/terminal-trajectories/ (matching run)`
20
+
21
+ ## Root cause
22
+
23
+ postActionVerifier labels a filesystem mtime observation outcomeClass='verified'. The no-todo runner path accepts that as direct task validation, and completionAutoFinalize can then report that changed files were validated. Observation of a successful mutation is weaker evidence than verification of requested behavior.
24
+
25
+ ## Code locations
26
+
27
+ - `packages/orchestrator/src/postActionVerifier.ts`
28
+ - `packages/orchestrator/src/taskValidationEvidence.ts`
29
+ - `packages/orchestrator/src/verificationInvocation.ts`
30
+ - `packages/orchestrator/src/verificationCommand.ts`
31
+ - `packages/orchestrator/src/agenticRunner.ts — no-todo validation recording and completion readiness`
32
+ - `packages/orchestrator/src/completionAutoFinalize.ts`
33
+ - `packages/orchestrator/tests/completionAutoFinalize.test.ts`
34
+ - `packages/orchestrator/tests/agenticRunner-completionReadiness.test.ts`
35
+ - `packages/orchestrator/tests/postActionVerifier.test.ts`
36
+
37
+ ## Repair design
38
+
39
+ - Represent mutation observation and task verification distinctly at the producer/consumer boundary.
40
+ - Prevent install/create/edit success and recent mtimes from minting task-validation authority.
41
+ - Retain explicit declared verifiers and real successful no-todo test/build verification, scoped to the actual execution and current mutation revision.
42
+ - Keep read-only/informational completion possible without inventing test obligations.
43
+ - Require fresh verification after a subsequent mutation and preserve truthful terminal summaries.
44
+
45
+ ## Acceptance checklist
46
+
47
+ - [x] Scaffold edits followed only by dependency installation do not auto-finalize as verified.
48
+ - [x] Mutation-only verifier observations cannot satisfy no-todo validation.
49
+ - [x] A declared verifier and a real successful no-todo test/build retain valid completion behavior.
50
+ - [x] Failed or stale verification cannot authorize completion; a later mutation invalidates earlier verification.
51
+ - [x] Read-only tasks do not require artificial build/test steps.
52
+ - [x] Focused completion/verifier tests and orchestrator typecheck pass.
53
+
54
+ ## Implementation and verification log
55
+
56
+ - Post-action audits now explicitly carry `completionAuthority: "none"`.
57
+ Their existing `verified` outcome remains an audit finding for telemetry;
58
+ it cannot establish acceptance of the user's task.
59
+ - Added typed task-validation evidence sourced only from a successful declared
60
+ verifier or an assertive validation execution. The selector excludes runtime
61
+ replays, contradictory/mismatched receipts, masked failures, and test/build
62
+ keywords used only as package arguments or file paths.
63
+ - No-todo automatic completion requires the typed evidence to match the
64
+ recorded command and turn. Tool-call identity invalidates a preceding pass
65
+ after another tool mutates files in the same turn, while preserving a build
66
+ that writes its own verified artifacts. Later failed checks revoke prior
67
+ validation even when no intervening mutation occurred.
68
+ - Added a mocked runner fixture with no subprocess or network execution. The
69
+ initial integration run reproduced both installation commands minting
70
+ `_lastBuildSuccessTurn=0`, and a subsequent same-turn edit leaving that
71
+ validation intact. These three assertions failed before runner integration.
72
+ - The explicit custom verifier fixture also exposed an existing producer gap:
73
+ `test -f ...` matched the declared verifier but the generic classifier
74
+ dismissed it as a read-only command. Its successful execution must reach the
75
+ completion ledger as declared verification; read-only tasks themselves gain
76
+ no new test obligation.
77
+ - Runner integration now routes explicit declared verifiers into the assertion
78
+ evidence ledger, carries typed authority into both automatic finalization
79
+ boundaries, and invalidates stale authority after mutations or failed checks.
80
+ - Verification completed September 5, 2026 UTC, from `packages/orchestrator`:
81
+ - `pnpm exec vitest run tests/taskValidationEvidence.test.ts tests/completionAutoFinalize.test.ts tests/postActionVerifier.test.ts tests/agenticRunner-taskValidationAuthority.test.ts tests/agenticRunner-completionReadiness.test.ts tests/verificationCommand.test.ts tests/completion-provenance-visual-evidence.test.ts` — **92 passed across 7 files**.
82
+ - `pnpm exec vitest run tests/agenticRunner.test.ts -t 'synthesizes truth-based completion when final-turn validation passes'` — **1 passed; 174 unrelated cases skipped**.
83
+ - `pnpm exec tsc --noEmit` — passed.
84
+ - The runner regressions use mock backends and mock shell tools. No live model
85
+ request, package installation, service mutation, or installed-package change
86
+ was performed. Git delivery is recorded by the parent work-order owner.
87
+
88
+ ### Independent-review followup after ce40bfca
89
+
90
+ - Review found that using the filesystem-audit intent parser also rejected real
91
+ commands with launcher options or executable paths, including `pnpm --filter
92
+ @omnius/cli build`, `npm --prefix app run build`, a local `.bin/vitest`, and
93
+ `npx --package typescript tsc --noEmit`. It also treated help/version and
94
+ collection modes as validation.
95
+ - Added a literal, quote-aware stage parser on the verification receipt
96
+ boundary and a separate invocation classifier. Launcher options consume
97
+ their values before action classification; executable basenames and inert
98
+ reporting suffixes are handled structurally. Unknown shell expansion stays
99
+ unclassified. Existing historical command matching remains unchanged.
100
+ - A matched declared command now has to verify the final state. For example,
101
+ `pnpm test && touch src/example.ts` cannot reuse the earlier test as proof of
102
+ the subsequent edit. `pnpm test && echo done` remains valid; redirected
103
+ reporting that writes a file does not count as an inert suffix.
104
+ - An untrusted later check or contradictory execution receipt revokes an
105
+ earlier pass even when the outer success flag is true. Positive protocol
106
+ text cannot override a nonzero exit receipt or restore that stale evidence.
107
+ - Final followup verification, from `packages/orchestrator`:
108
+ - `pnpm exec vitest run tests/taskValidationEvidence.test.ts tests/verificationInvocation.test.ts tests/completionAutoFinalize.test.ts tests/postActionVerifier.test.ts tests/agenticRunner-taskValidationAuthority.test.ts tests/agenticRunner-completionReadiness.test.ts tests/verificationCommand.test.ts tests/completion-provenance-visual-evidence.test.ts` — **168 passed across 8 files**, including **25 runner cases**.
109
+ - `pnpm exec tsc --noEmit` — passed.
110
+ - All review examples have deterministic regressions. Followup Git delivery is
111
+ pending the parent owner's scoped commit; no runtime publication is implied.
112
+
113
+ ### Full-suite compatibility followup
114
+
115
+ - The stale-review reconciliation fixture depended on the removed generic
116
+ completion heuristic despite declaring no verifier or edited artifact. It
117
+ now supplies a completed todo with `pnpm test` and a coherent execution
118
+ receipt, asserts declared verification authority, and disables automatic
119
+ finalization only to reach the explicit terminal boundary under test.
120
+ - Independent review also reproduced rejected valid Vitest `--run` and
121
+ `--coverage` invocations. Vitest mode options now use their own arity table,
122
+ separate from package-launcher options. Collection still remains inspection
123
+ when mode flags precede `list`; a config file named `list` remains a value.
124
+ - Verification from `packages/orchestrator`:
125
+ - `pnpm exec vitest run tests/verificationInvocation.test.ts tests/run-observability.test.ts tests/agenticRunner-taskValidationAuthority.test.ts` — **93 passed across 3 files** (57 invocation, 10 observability, 26 runner cases).
126
+ - `pnpm exec tsc --noEmit` — passed.
127
+ - This followup uses mocked execution and adds no live inference or runtime
128
+ state changes. Parent work-order owner handles the scoped commit and push.
129
+
130
+ ## Delivery
131
+
132
+ - Commits: `ce40bfca, cf966ea3, dc60e5bf` on `origin/main`.
133
+ - Clean `pnpm -r build` passed for all 11 workspace packages/apps. Final aggregate suite evidence is in the linked program ledger.
134
+
135
+ ## Publication boundary
136
+
137
+ Repository repair and deterministic verification are this order's completion
138
+ boundary. The user will publish the package and then perform live validation.
139
+ No package publication, installed-package replacement, runtime-state repair,
140
+ or service restart is authorized by this order.
@@ -0,0 +1,82 @@
1
+ # WO-28: Retire obsolete task evidence after explicit scope changes
2
+
3
+ **Status:** complete; implemented, verified, and pushed to origin/main
4
+ **Priority:** P2
5
+ **Incident date:** 2026-09-04 PDT / 2026-09-05 UTC
6
+ **Program:** [September 4 Telegram long-haul repairs](TELEGRAM-LONGHAUL-2026-09-04.md)
7
+
8
+ ## Observed failure
9
+
10
+ At 20:41 PDT, current turn 21 retained 57,931 characters of previous pentest repository tool bodies (38.6% of all message content), versus 35,088 from the new boutique-services project. The model eventually followed the new focus, but old task material remained in the active frame/history and sustained raw-discovery pressure.
11
+
12
+ ## Evidence
13
+
14
+ Paths are relative to /home/roko/Documents/Projects/Adjacent. Live files are
15
+ observational references; sanitized deterministic fixtures must reproduce the
16
+ failure without copying credentials or private conversation content.
17
+
18
+ - `telegram_test/.omnius/context-window-dumps/2026-09-05T03-41-09-548Z-main-57e9eeba54.json`
19
+
20
+ ## Root cause
21
+
22
+ Telegram deliberately admits normal messages as `context_only`. The typed model reconciliation had no replacement field, and an immediately following voice message overwrote the one unresolved steering slot. Even explicit transport replacements reset only part of the task state: raw messages, evidence-ledger bodies, tool-event context, and task trajectory survived. Recovery also retained an old lifecycle epoch.
23
+
24
+ ## Code locations
25
+
26
+ - `packages/orchestrator/src/agenticRunner.ts`: grouped steering, typed replacement, exact archives, safe checkpoints, task-state reset, request/recovery integration.
27
+ - `packages/orchestrator/src/typed-model-output.ts` and `steeringIntake.ts`: bounded typed `taskDisposition`, `replacementGoal`, retained tool IDs.
28
+ - `packages/orchestrator/src/taskScopeBoundary.ts`: validated run/session/epoch selector and complete tool transactions.
29
+ - `packages/orchestrator/src/interruption-lifecycle.ts`: drained, owner-checked epoch advancement with persistence rollback.
30
+ - `packages/orchestrator/src/contentAddressedArtifactStore.ts`: explicit durable flush for transition archives.
31
+ - `packages/orchestrator/tests/agenticRunner-task-scope.test.ts`, `taskScopeBoundary.test.ts`, `typed-model-output.test.ts`, lifecycle/artifact tests and production request replay.
32
+
33
+ ## Repair design
34
+
35
+ - Trace and repair the typed task-replacement boundary; do not infer replacement from filenames or keyword similarity alone.
36
+ - Retire prior-task raw tool transactions from the next model-visible request after an accepted explicit replacement, preserving their durable audit/receipt references.
37
+ - Retain current user authority, replacement instructions, media reference evidence needed to interpret the new task, and explicitly adopted cross-task dependencies.
38
+ - Do not treat a status question, additive instruction, or priority change as destructive task replacement.
39
+ - Keep canonical goal and evidence selection consistent across continuation/restoration and integrate with WO-25.
40
+
41
+ ## Acceptance checklist
42
+
43
+ - [x] A two-project replacement replay excludes obsolete raw source while preserving current task instructions and required new evidence.
44
+ - [x] Status-only and additive steering preserve relevant active-task context.
45
+ - [x] Retired evidence remains durably retrievable; tool call/result pairs remain valid.
46
+ - [x] Restored continuation uses the same scope boundary without resurrecting obsolete raw history.
47
+ - [x] Integrated compiler/projection replay proves scope retirement does not trigger useless compaction.
48
+ - [x] Focused steering/context tests and orchestrator typecheck pass.
49
+
50
+ ## Implementation and verification log
51
+
52
+ - Initial regression reproduced overwritten text authority, retained obsolete raw history, and unvalidated retained tool IDs (three failing cases).
53
+ - Typed reconciliation batches unresolved text/media and advances scope only for an accepted replacement, with no lexical task classifier. Transport replacement acknowledgement cannot advance twice.
54
+ - Exact immutable transition archives include messages, todos, workboard, completion ledger, and task state. File/directory sync precedes lifecycle epoch advancement and retirement.
55
+ - Later safe transcripts are checkpointed separately so restart retains work performed after replacement. Early recovery selects current goal before startup analysis.
56
+ - Reset covers secondary evidence/context caches, world facts, task trajectory, phase tree, and verifier authority. Explicitly adopted complete tool transactions remain model-visible.
57
+ - Independent review exposed and drove repairs for cache reinjection, structured-state archive loss, restore progress loss, and identical text in later task scope. Final verification and commit evidence follow below.
58
+
59
+ ## Verification result
60
+
61
+ - Runner scope regressions: **13 passed**; pure boundary regressions: **11 passed**.
62
+ - Durable interruption lifecycle: **25 passed**; content-addressed archive store: **10 passed**.
63
+ - Full mocked `runner.run()` replacement replay: **2 passed**, exercising primary and brute-force request paths, grouped text/audio, exact epoch advancement, evidence retirement, and retained core policy.
64
+ - Existing core-policy regressions: **3 passed**, covering small, medium, and large tiers.
65
+ - Compiler/canonical/steering integration regression set passed. Clean workspace build passed for all 11 packages/apps.
66
+ - Final scope reset includes observation and source ledgers, contract registry, local tool cache, world facts, recovery-controller delta, phase tree, and cached prior-task retrieval. Current workspace inventory remains observable.
67
+ - Identical later user replies survive restore: selectors recognize the original archived prefix and stop retirement at the new authority/epoch boundary.
68
+ - Runtime policy has typed metadata and stays in separate bounded system messages. This avoids losing core rules to first-system truncation while retiring old task recall.
69
+ - Full aggregate suite and commit evidence are recorded in the program ledger.
70
+
71
+ ## Delivery
72
+
73
+ - Commit: `24321b30` on `origin/main`.
74
+ - Final archive-directory sync refinement passed both scope suites (**15 tests**) and an additional workspace build. Full aggregate evidence is in the linked program ledger.
75
+
76
+ ## Publication boundary
77
+
78
+ Repository repair and deterministic verification are this order's completion
79
+ boundary. The user will publish the package and then perform live validation.
80
+ No package publication, installed-package replacement, runtime-state repair,
81
+ or service restart is authorized by this order.
82
+
@@ -0,0 +1,85 @@
1
+ # WO-29: Exclude retired failures from active diagnostics
2
+
3
+ **Status:** complete; implemented, verified, and pushed to origin/main
4
+ **Priority:** P2
5
+ **Incident date:** 2026-09-04 PDT / 2026-09-05 UTC
6
+ **Program:** [September 4 Telegram long-haul repairs](TELEGRAM-LONGHAUL-2026-09-04.md)
7
+
8
+ ## Observed failure
9
+
10
+ Current transcription and shell-timeout recovery cards were retired at 19:58 PDT with supersededBy='failure-expired:turn-19', yet unresolved-failure diagnostics still appeared at 20:35. This is contradictory active recovery context; it was not proven to be a terminal blocker.
11
+
12
+ ## Evidence
13
+
14
+ Paths are relative to /home/roko/Documents/Projects/Adjacent. Live files are
15
+ observational references; sanitized deterministic fixtures must reproduce the
16
+ failure without copying credentials or private conversation content.
17
+
18
+ - `telegram_test/.omnius/workboards/telegram-64ac9937dc7647a8-1788572475259-1/active.json`
19
+
20
+ ## Root cause
21
+
22
+ Failure retirement preserves status for audit and marks `supersededBy`.
23
+ Workboard active diagnostic derivation checked failure card ID/status but
24
+ ignored supersession. The same omission affected compact card counts and
25
+ selection, synthesis blockers and verified evidence, and the diagnostics tool
26
+ view. Cached snapshots could retain the old derived warnings even after a
27
+ runtime upgrade.
28
+
29
+ ## Code locations
30
+
31
+ - `packages/execution/src/tools/workboard.ts — active-card selection, diagnostic derivation, projections`
32
+ - `packages/execution/tests/workboard.test.ts`
33
+ - `packages/orchestrator/tests/workboard-run-continuity.test.ts`
34
+
35
+ ## Repair design
36
+
37
+ - Use consistent active-card lifecycle semantics in derived workboard diagnostics and active projections.
38
+ - Exclude superseded/expired obligations from unresolved counts and actionable warnings while preserving historical cards and diagnostic records where explicitly historical.
39
+ - Do not relabel expiration as successful verification.
40
+ - Ensure a genuinely new recurrence creates/activates the correct obligation without resurrecting retired state accidentally.
41
+
42
+ ## Acceptance checklist
43
+
44
+ - [x] An expired/superseded in-progress failure no longer emits active unresolved-failure diagnostics.
45
+ - [x] A current unresolved failure still emits its bounded warning.
46
+ - [x] Retirement preserves the original failure and its audit trail without fabricating verification.
47
+ - [x] Fresh recurrence is represented correctly and legacy cards without supersession retain behavior.
48
+ - [x] Focused workboard/lifecycle tests and execution typecheck pass.
49
+
50
+ ## Implementation and verification log
51
+
52
+ - A single active-card predicate now excludes `supersededBy` cards from compact
53
+ counts, card limits, hidden-card counts, synthesis, and newly derived failure
54
+ and dependency diagnostics. The existing ownership-frontier check uses the
55
+ same predicate.
56
+ - Active diagnostic projection filters recorded or cached historical warnings
57
+ for retired cards before model-output limits are applied. Full JSON, raw
58
+ cards, status, evidence, and append-only events preserve the audit history.
59
+ An explicitly returned compact card includes its supersession marker.
60
+ - Retirement supplies no verification evidence. A live card depending on a
61
+ retired but unverified card remains blocked; the dependent obligation must
62
+ be explicitly reconciled. A new recurrence receives its own active warning.
63
+ - Before the repair, regression fixtures reproduced retired-card
64
+ `unresolved_failure_card`/`repeated_worker_failure` warnings and a board
65
+ containing only retired cards failing to report its empty active frontier.
66
+ - Hermetic verification on 2026-09-05 UTC:
67
+ - `pnpm --filter @omnius/execution exec vitest run tests/workboard.test.ts`:
68
+ **37 tests passed**, including cached-snapshot projections, replay/reload,
69
+ preserved failed evidence, fresh recurrence, and dependency safety.
70
+ - `pnpm --filter @omnius/orchestrator exec vitest run tests/workboard-run-continuity.test.ts`:
71
+ **10 tests passed**.
72
+ - `pnpm --filter @omnius/execution typecheck`: **passed**.
73
+ - Repository delivery commit will be recorded by the parent repair task.
74
+
75
+ ## Delivery
76
+
77
+ - Commits: `3a1731ad` on `origin/main`.
78
+ - Clean `pnpm -r build` passed for all 11 workspace packages/apps. Final aggregate suite evidence is in the linked program ledger.
79
+
80
+ ## Publication boundary
81
+
82
+ Repository repair and deterministic verification are this order's completion
83
+ boundary. The user will publish the package and then perform live validation.
84
+ No package publication, installed-package replacement, runtime-state repair,
85
+ or service restart is authorized by this order.
@@ -1,12 +1,12 @@
1
1
  {
2
2
  "name": "omnius",
3
- "version": "1.0.694",
3
+ "version": "1.0.695",
4
4
  "lockfileVersion": 3,
5
5
  "requires": true,
6
6
  "packages": {
7
7
  "": {
8
8
  "name": "omnius",
9
- "version": "1.0.694",
9
+ "version": "1.0.695",
10
10
  "bundleDependencies": [
11
11
  "image-to-ascii"
12
12
  ],
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "omnius",
3
- "version": "1.0.694",
3
+ "version": "1.0.695",
4
4
  "description": "AI coding agent powered by open-source models (Ollama/vLLM) — interactive TUI with agentic tool-calling loop",
5
5
  "type": "module",
6
6
  "main": "./dist/library.js",