omnius 1.0.689 → 1.0.691

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -33824,6 +33824,84 @@
33824
33824
  }
33825
33825
  ]
33826
33826
  },
33827
+ {
33828
+ "id": "guide.work-orders-runtime-health-remediation-wo-21-task-convergence-and-todo-scope-uppercase",
33829
+ "kind": "guide",
33830
+ "title": "WO-21: Task convergence and todo scope",
33831
+ "summary": "On 2026-09-03, a user reported that Omnius remained at task 2/8 for about 40 turns. The report did not include a screenshot or run artifact, so the exact label remains unverified. Omnius has two similar counters:",
33832
+ "keywords": [
33833
+ "work",
33834
+ "orders",
33835
+ "runtime",
33836
+ "health",
33837
+ "remediation",
33838
+ "WO",
33839
+ "21",
33840
+ "task",
33841
+ "convergence",
33842
+ "and",
33843
+ "todo",
33844
+ "scope",
33845
+ "md"
33846
+ ],
33847
+ "maturity": "internal",
33848
+ "audiences": [
33849
+ "maintainer",
33850
+ "large-context-agent"
33851
+ ],
33852
+ "layer": "documentation",
33853
+ "interfaces": [
33854
+ {
33855
+ "type": "file",
33856
+ "target": "docs/work-orders/runtime-health-remediation/WO-21-task-convergence-and-todo-scope.md"
33857
+ }
33858
+ ],
33859
+ "references": [
33860
+ {
33861
+ "type": "documentation",
33862
+ "target": "docs/work-orders/runtime-health-remediation/WO-21-task-convergence-and-todo-scope.md",
33863
+ "relation": "canonical-artifact"
33864
+ }
33865
+ ]
33866
+ },
33867
+ {
33868
+ "id": "guide.work-orders-telegram-dmn-wo-22-dmn-outreach-and-learning-uppercase",
33869
+ "kind": "guide",
33870
+ "title": "WO-22: DMN outreach, DM sharing, and outcome learning",
33871
+ "summary": "On 2026-09-03 at 17:07 PDT the bot posted in the OMNIUS group without being addressed. The operator asked whether this was self-induced reflection.",
33872
+ "keywords": [
33873
+ "work",
33874
+ "orders",
33875
+ "telegram",
33876
+ "dmn",
33877
+ "WO",
33878
+ "22",
33879
+ "dmn",
33880
+ "outreach",
33881
+ "and",
33882
+ "learning",
33883
+ "md"
33884
+ ],
33885
+ "maturity": "internal",
33886
+ "audiences": [
33887
+ "maintainer",
33888
+ "large-context-agent"
33889
+ ],
33890
+ "layer": "documentation",
33891
+ "interfaces": [
33892
+ {
33893
+ "type": "file",
33894
+ "target": "docs/work-orders/telegram-dmn/WO-22-dmn-outreach-and-learning.md"
33895
+ }
33896
+ ],
33897
+ "references": [
33898
+ {
33899
+ "type": "documentation",
33900
+ "target": "docs/work-orders/telegram-dmn/WO-22-dmn-outreach-and-learning.md",
33901
+ "relation": "canonical-artifact"
33902
+ }
33903
+ ]
33904
+ },
33827
33905
  {
33828
33906
  "id": "guide.work-orders-telegram-dropbear-context-rca-workorder",
33829
33907
  "kind": "guide",
package/docs/DISCOVERY.md CHANGED
@@ -559,6 +559,8 @@ Daemon equivalents are `GET /v1/discovery/bootstrap`, `GET /v1/discovery?q=<inte
559
559
  | `guide.work-orders-runtime-health-remediation-wo-19-clean-build-evidence-uppercase` | WO-19 Clean Build Evidence | WO-19 is implemented in commit 5f253c05. |
560
560
  | `guide.work-orders-runtime-health-remediation-wo-19-clean-build-reproducibility-uppercase` | WO-19: Clean Build and Optional-Dependency Reproducibility | Status: deterministic acceptance complete Risk: medium Depends on: WO-00 |
561
561
  | `guide.work-orders-runtime-health-remediation-wo-20-exact-session-continuation-uppercase` | WO-20: Exact session continuation across exit and restart | On 2026-09-03, the telegramtest TUI displayed the latest pentest task during startup, but the first restored model context combined a completed breach.py handoff with an unrelated, older page.tsx todo tree. The user then entered continue, and Omnius continued the stale tree instead of the task that was active before /quit. |
562
+ | `guide.work-orders-runtime-health-remediation-wo-21-task-convergence-and-todo-scope-uppercase` | WO-21: Task convergence and todo scope | On 2026-09-03, a user reported that Omnius remained at task 2/8 for about 40 turns. The report did not include a screenshot or run artifact, so the exact label remains unverified. Omnius has two similar counters: |
563
+ | `guide.work-orders-telegram-dmn-wo-22-dmn-outreach-and-learning-uppercase` | WO-22: DMN outreach, DM sharing, and outcome learning | On 2026-09-03 at 17:07 PDT the bot posted in the OMNIUS group without being addressed. The operator asked whether this was self-induced reflection. |
562
564
  | `guide.work-orders-telegram-dropbear-context-rca-workorder` | Telegram Dropbear Context Engineering RCA Work Order | Observed run: /home/roko/Documents/Projects/Adjacent/telegramtest/.omnius, run id 1782873796963-i5r7mv. |
563
565
  | `guide.work-orders-wo-am-gaps-uppercase` | Associative Memory Gap Work Orders | Generated: 2026-04-13 Source: Deep audit of multimodal associative memory systems Status: READY FOR IMPLEMENTATION |
564
566
  | `guide.work-orders-world-class-memory-compiler-readme-uppercase` | World-Class Memory Compiler Program | Status: active implementation program Owner: Omnius orchestration and memory packages Last updated: 2026-07-13 |
@@ -311,3 +311,15 @@ deterministically repaired P0 or P1 defect.
311
311
  pass cancellation, restart, ownership, external-effect, and isolation tests.
312
312
  - [x] WO-20 exit and restart preserve one exact typed task continuation. Generic
313
313
  context restoration cannot mix or activate unrelated historical artifacts.
314
+ - [x] WO-21 task convergence and todo scope prevent stale `tasks N/M` displays,
315
+ cross-runner todo leakage, and activity-only re-engagement.
316
+ - [x] Bind todo tools to immutable runner-owned sessions.
317
+ - [x] Clear the active persisted checklist at a fresh task boundary after
318
+ archival.
319
+ - [x] Count only typed authoritative advancement as progress.
320
+ - [x] Latch an advisory convergence review when the leaf frontier is static.
321
+ - [x] Make loop, reminder, budget, and workboard progress semantics truthful.
322
+ - [x] Pass deterministic race, stasis, display, typecheck, and build tests.
323
+ - [x] A recovered pause no longer strands its session. Submitting a new task
324
+ retires the superseded generation instead of refusing every later prompt,
325
+ while unreconciled external effects still fence admission.
@@ -0,0 +1,115 @@
1
+ # WO-21: Task convergence and todo scope
2
+
3
+ ## Incident
4
+
5
+ On 2026-09-03, a user reported that Omnius remained at `task 2/8` for about
6
+ 40 turns. The report did not include a screenshot or run artifact, so the
7
+ exact label remains unverified. Omnius has two similar counters:
8
+
9
+ - `tasks 2/8` is the completed-leaf count in the pinned TUI checklist.
10
+ - `Loop intervention 2/8` is the second repetition intervention in a
11
+ model-tier-specific series.
12
+
13
+ The source audit found defects in both paths. These defects are sufficient to
14
+ produce the reported symptom even though they do not prove which counter the
15
+ reporter saw.
16
+
17
+ ## Root causes
18
+
19
+ - A TUI process reuses one todo session across tasks. A fresh runner hides old
20
+ todo IDs in its private view but leaves the persisted checklist unchanged.
21
+ The TUI reads that unchanged file and can display a stale `tasks 2/8` row.
22
+ - Todo tools fall back to one mutable process-global session ID. Concurrent
23
+ parent, Telegram, background, and child runners can redirect one another's
24
+ unscoped reads and writes.
25
+ - Any unique non-noop tool result increments the counter used to permit turn
26
+ extension and brute-force re-engagement. Varying reads, searches, and error
27
+ text therefore masquerade as task advancement.
28
+ - The repeated-loop counter resets after its maximum intervention even though
29
+ no task state changed. Its status says that user guidance was requested, but
30
+ it does not create a user-input boundary.
31
+ - A failed `todo_write` advances the reminder clock. The visible checklist can
32
+ remain unchanged for ten more turns without a reminder.
33
+ - Small-model budget text says that a todo update resets the phase budget, but
34
+ the implementation resets only after a context-tree phase transition.
35
+ - Workboard `lastSubstantiveProgress` treats successful discovery reads as
36
+ progress. This conflicts with the task and completion ledgers, where a read
37
+ is evidence but not advancement.
38
+
39
+ ## Required invariants
40
+
41
+ - [x] Every runner binds todo tools to its own immutable session scope.
42
+ - [x] Model-supplied todo session IDs cannot override the host scope.
43
+ - [x] Concurrent runners cannot redirect one another's todo reads or writes.
44
+ - [x] A fresh task archives the prior checklist and clears the active persisted
45
+ projection so the TUI, runner, REST, and Telegram views agree.
46
+ - [x] Only a confirmed mutation, todo transition, workboard transition,
47
+ assertive verifier receipt, or delivery receipt counts as authoritative task
48
+ progress.
49
+ - [x] Reads, searches, runtime-authored blocks, no-ops, and changing errors do
50
+ not extend the run or re-arm brute-force execution. Entry into the first
51
+ re-engagement cycle is deliberately not gated on prior authoritative
52
+ progress: that cycle is the rescue attempt for a run that has only read.
53
+ If it also advances nothing, re-engagement stops at cycle 2.
54
+ - [x] Repeated stasis creates one latched, model-visible convergence review.
55
+ - [x] The convergence review remains advisory. It does not infer completion,
56
+ manufacture a blocker, or force a user question.
57
+ - [x] The loop intervention maximum does not silently reset without
58
+ authoritative progress.
59
+ - [x] Todo reminders advance only after a successful state-changing write.
60
+ - [x] Budget exhaustion text describes the actual reset condition.
61
+ - [x] Workboard progress metadata uses the same authoritative distinction.
62
+ - [x] Todo stagnation compares completed leaves, matching the TUI counter.
63
+
64
+ ## Status
65
+
66
+ Implemented and verified on 2026-09-03.
67
+
68
+ | Invariant | Implementation | Test |
69
+ | --- | --- | --- |
70
+ | Immutable runner-owned todo scope | `packages/execution/src/tools/todo-write.ts` `bindExecutionScope` | `todo-store.test.ts` "binds todo tools to a host session that model arguments cannot replace" |
71
+ | Concurrent runner isolation | `AgenticRunner` constructor resolves one immutable `_sessionId`; `registerTool` binds it | `todo-store.test.ts` "keeps concurrent runner-scoped todo writes isolated" |
72
+ | Fresh task clears the active projection | `_hidePriorSessionTodosForFreshTask` writes an empty active checklist after archival | `fresh-task-todo-boundary.test.ts` |
73
+ | Typed authoritative progress only | `packages/orchestrator/src/convergence-progress.ts` | `task-convergence-progress.test.ts` (8 cases) |
74
+ | Latched advisory convergence review | `agenticRunner.ts` convergence review block | `task-convergence-progress.test.ts` "latches one advisory review and clears it only after typed progress" |
75
+ | Loop intervention maximum stays latched | `Math.min(maxInterventions, loopInterventionCount + 1)` | `agenticRunner-context-behavior.test.ts` "latches the loop-intervention maximum instead of cycling it back to one" |
76
+ | Workboard progress uses the same distinction | `recordWorkboardToolCall` returns a status-fingerprint transition | `workboard-run-continuity.test.ts` "does not record a successful discovery read as task advancement" |
77
+
78
+ Full suites run green: `@omnius/execution` 1507 passed, `@omnius/orchestrator`
79
+ 2303 passed (200 files, 2 skipped), `omnius` CLI 2428 passed (258 files).
80
+ Typecheck passes for all three packages and `pnpm -r build` completes.
81
+
82
+ The reported counter was never attributed to a specific field run. The evidence
83
+ request below stays open.
84
+
85
+ One correction was made during verification. An earlier draft of this work also
86
+ refused to enter brute-force cycle 1 whenever the primary run recorded no
87
+ authoritative advancement. That inverted the purpose of re-engagement and broke
88
+ eight existing brute-force tests, because a run that has only read is exactly
89
+ the run that still needs one push to act. The cycle-2 gate already bounds the
90
+ thrash this work order set out to stop, so the cycle-1 refusal was removed.
91
+
92
+ ## Deterministic verification
93
+
94
+ - [x] Scoped todo tools ignore conflicting model-supplied session IDs.
95
+ - [x] Concurrent scoped todo tools write isolated session files.
96
+ - [x] A same-session fresh task clears the old visible todo projection after
97
+ archiving it.
98
+ - [x] Forty distinct reads produce zero authoritative progress.
99
+ - [x] Changing failures produce zero authoritative progress.
100
+ - [x] Mutations, todo transitions, workboard transitions, assertive verifiers,
101
+ and delivery receipts produce typed progress.
102
+ - [x] A convergence review latches after the configured unchanged frontier and
103
+ clears only after typed progress.
104
+ - [x] The loop intervention maximum stays latched instead of cycling to one.
105
+ - [x] Focused execution, orchestrator, and CLI tests pass.
106
+ - [x] Affected package typechecks and builds pass.
107
+
108
+ ## Evidence request for exact field attribution
109
+
110
+ If the original run is available, collect only a redacted structural slice:
111
+ the exact counter line or screenshot, Omnius version, model/backend, launch
112
+ surface, incident time and timezone, session/run IDs, five turns before the
113
+ stall through ten turns after it, todo IDs/statuses/revisions, workboard card
114
+ statuses, and completion-ledger statuses. Do not copy a complete `.omnius`
115
+ directory because debug previews can contain user text and source fragments.
@@ -0,0 +1,85 @@
1
+ # WO-22: DMN outreach, DM sharing, and outcome learning
2
+
3
+ ## Incident
4
+
5
+ On 2026-09-03 at 17:07 PDT the bot posted in the OMNIUS group without being
6
+ addressed. The operator asked whether this was self-induced reflection.
7
+
8
+ It was not. The post was reactive to two inbound messages (5680, 5681) and
9
+ authorized as `direct_turn`. The stated router reason — *"Sender reacts to
10
+ bot's recent output in active thread"* — was false: the cited self-bearing
11
+ evidence was `turn:5678`, a bot reply from **09-01**, to a different user,
12
+ about a different topic.
13
+
14
+ Two independent defects produced that:
15
+
16
+ - `buildTelegramInteractionEvidencePacket` selects causal history with
17
+ `.slice(-8)` and **no time bound**. In a quiet room the last 8 messages
18
+ reach back days.
19
+ - The `direct_turn` validator requires a cited self-bearing evidence entry but
20
+ checks only *kind* and *actor*, never *recency*. A two-day-old turn satisfied
21
+ it.
22
+
23
+ Investigating that surfaced the larger finding below.
24
+
25
+ ## The real finding: reflection could not speak
26
+
27
+ Omnius reflects constantly — **465 daydream artifacts on 2026-09-03 alone** —
28
+ and roughly two thirds of recent artifacts proposed a same-group follow-up.
29
+ None were ever sent. Across 639 committed routing decisions in this workspace,
30
+ `social_intervention` fired exactly **once**; the other 432 replies were all
31
+ `direct_turn`.
32
+
33
+ Root cause was ID plumbing, not judgement:
34
+
35
+ - The extraction schema advertised `reply_to_message_id: 1` and
36
+ `source_message_ids: [1]` as examples. **78.6% of live follow-up proposals
37
+ echoed a placeholder-like id**, `1` being the single most common value.
38
+ - Those ids never matched real Telegram ids, so `candidateMessageIds` was
39
+ almost always empty.
40
+ - The discretion prompt then instructed the model to *"cite at least one
41
+ current eligible Telegram message ID"* **without ever listing which ids were
42
+ eligible**, while the host discarded any follow-up citing an id outside that
43
+ undisclosed set.
44
+
45
+ Silence was the only reachable outcome.
46
+
47
+ ## Changes
48
+
49
+ | Area | Change |
50
+ | --- | --- |
51
+ | Extraction prompt | Schema example ids replaced with non-numeric sentinels; explicit rule that ids must appear literally in the corpus; the valid id set is now stated. |
52
+ | Extraction result | `runTelegramReflectionExtraction` drops any `source_message_ids` / `reply_to_message_id` absent from the corpus, so a copied or hallucinated anchor cannot reach the follow-up gate. |
53
+ | Discretion prompt | Now supplies `Eligible evidence message IDs`, the exact set the host validates against. |
54
+ | DM reflection | Idle reflection and outreach admitted for private chats **only after that chat opts in** via `/reflect auto on`. Default remains off. |
55
+ | Outcome learning | New `telegram-outreach-outcomes.ts` records every autonomous send and settles it as engaged / ignored from subsequent room traffic. Settled history is fed back into the next discretion call. |
56
+
57
+ ## Design notes
58
+
59
+ - **DMs are opt-in, not default.** A DM is one person's inbox. The capability
60
+ exists; the host does not enable it on the recipient's behalf.
61
+ - **Silence is recorded as a real signal.** An unanswered outreach settles to
62
+ `engaged: false` after 45 minutes rather than staying unknown, so the
63
+ calibration block cannot mistake absence of data for success.
64
+ - **The bot's own traffic never counts as engagement.** Only a human message
65
+ inside the window closes an outcome positively.
66
+ - **Calibration is neutral before it is negative.** With no settled history the
67
+ prompt says so and asks the model to judge on merit; it raises the bar only
68
+ once every settled attempt was ignored.
69
+
70
+ ## Not addressed
71
+
72
+ The evidence-packet recency defect that produced the original false
73
+ `direct_turn` remains open. Bounding causal history by time would push honest
74
+ unaddressed interjections onto the `social_intervention` path, where they are
75
+ visible and tunable — but `social_intervention` currently has **no validation
76
+ rules at all** (compare five checks for `direct_turn`). Both should be done
77
+ together so interjection becomes a governed capability rather than an
78
+ unvalidated fallback.
79
+
80
+ ## Verification
81
+
82
+ CLI suite: 2442 passed, 0 failed (261 files). Typecheck clean; `pnpm -r build`
83
+ clean. 14 new tests covering id anchoring, corpus allow-list enforcement,
84
+ outcome settlement, calibration wording, torn-ledger recovery, and bridge
85
+ wiring.
@@ -1,12 +1,12 @@
1
1
  {
2
2
  "name": "omnius",
3
- "version": "1.0.689",
3
+ "version": "1.0.691",
4
4
  "lockfileVersion": 3,
5
5
  "requires": true,
6
6
  "packages": {
7
7
  "": {
8
8
  "name": "omnius",
9
- "version": "1.0.689",
9
+ "version": "1.0.691",
10
10
  "bundleDependencies": [
11
11
  "image-to-ascii"
12
12
  ],
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "omnius",
3
- "version": "1.0.689",
3
+ "version": "1.0.691",
4
4
  "description": "AI coding agent powered by open-source models (Ollama/vLLM) — interactive TUI with agentic tool-calling loop",
5
5
  "type": "module",
6
6
  "main": "./dist/library.js",