omnius 1.0.696 → 1.0.698

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,42 @@
1
+ # WO-46: Preserve process authority for shell observations
2
+
3
+ Status: repository repair complete and delivered to `origin/main`. Publication and live validation remain with the user.
4
+
5
+ ## Finding and root repair
6
+
7
+ Field review F1 recorded successful source reads changed into failures because source text contained `throw new Error`. The runner's semantic stdout scanner and ShellTool's own soft-failure scanner both infer process failure from arbitrary output. A controlled reproduction also finds a false failure after `cd . && cat` prints a documented error line.
8
+
9
+ Preserve observed command/process outcomes through both execution loops. Reconcile actual nonzero exits, timeout, primary-command failures and properly bound verifier protocol results before workflow settlement. Source or documentation reads must not manufacture runtime or verification failures from their bytes. Share existing command semantics where needed to bind explicit verifier output; retain existing import compatibility.
10
+
11
+ ## Owned paths
12
+
13
+ - `packages/orchestrator/src/agenticRunner.ts`: shared tool result reconciliation and removal of generic semantic failure authority.
14
+ - `packages/execution/src/tools/shell.ts`: process/result producer and explicit verifier binding.
15
+ - Existing command-classification modules and compatibility exports, if required for one shared interpretation.
16
+ - Shell and actual-runner regression tests for both loops.
17
+
18
+ ## Acceptance
19
+
20
+ - [x] Successful source reads containing exception, compiler, test and verifier-marker literals remain successful observations.
21
+ - [x] Real nonzero, timeout, masked pipeline primary failures and bound explicit verification failures remain failed.
22
+ - [x] Both production loops and workflow settlement observe the same typed result.
23
+ - [x] No source output is promoted into verification authority; existing task/shell regressions pass.
24
+ - [x] Record full package verification, scoped commit and push.
25
+
26
+ ## Verification evidence
27
+
28
+ - 192 focused orchestrator tests passed, including both execution loops, actual ShellTool reads and verifier controls. Log: `/tmp/omnius-wo46-focused-final.log`.
29
+ - 63 focused execution tests passed; the complete execution package then passed 1,820 tests with 3 pre-existing skips across 163 suites. Log: `/tmp/omnius-remedies-execution-full.log`.
30
+ - Execution build, orchestrator typecheck and whitespace checks passed.
31
+ - Terminal-report handling of accepted negative searches and preview exits is completed with WO48, using the same command classification. Publication/live validation remain pending with the user.
32
+
33
+ ## Repository acceptance and delivery
34
+
35
+ - Source repair: `a3c26eca`, delivered to `origin/main`; coordinated changes: `672bef53 (terminal observation/report integration)`.
36
+ - Clean rebuild of all workspace packages and final orchestrator rebuild passed.
37
+ - Execution: 1,820 passed / 3 existing skips across 163 suites.
38
+ - CLI: 2,575 passed across 274 suites. Rerun with `--maxWorkers=4 --minWorkers=1` cleared the original five UI import timeouts without changing those tests or their limits.
39
+ - Orchestrator: full run covered 228 suites with 2,847 passed / 1 existing skip / 1 outdated fixture failure. The fixture lacked required `pipefail`; it was corrected, and all 10 tests in that suite then passed. No production source changed after the full run began. This is full-suite coverage plus the corrected suite rerun, not a claim that the initial command exited successfully.
40
+ - Across these packages, all 7,243 distinct non-skipped tests have a passing result on the final implementation. Logs: `/tmp/omnius-remedies-execution-full.log`, `/tmp/omnius-remedies-orchestrator-full.log`, `/tmp/omnius-remedies-receipt-fixture.log`, `/tmp/omnius-remedies-cli-final.log`.
41
+
42
+ Checks used hermetic backend/Telegram transport and synthetic local process/project fixtures. No live inference, service restart, installed-package replacement, Telegram sends or publication were performed. The existing publish staging and unrelated discovery changes were preserved.
@@ -0,0 +1,54 @@
1
+ # WO-47: Validate and execute diagnostics without blocking the agent
2
+
3
+ Status: repository repair complete and delivered to `origin/main`. Publication and live validation remain with the user.
4
+
5
+ ## Finding and root repair
6
+
7
+ Field review F2 recorded `requestedSteps.filter is not a function`: a JSON-encoded string reached an unchecked array cast. Review also found silently omitted requested checks, successful 0/0 results, implicit package downloads, and synchronous subprocess calls that can block typing and Stop for two minutes per step.
8
+
9
+ Use an executable input contract at runner admission and direct invocation. Validate requested steps and project paths before launching anything. Report unavailable requested checks and empty detection explicitly. Use configured local commands through the existing asynchronous process runner, with bounded output, timeout and cancellation of the owning process group. Preserve actual per-step outcomes in typed command receipts; never invent an aggregate successful command.
10
+
11
+ ## Owned paths
12
+
13
+ - `packages/execution/src/tools/diagnostic.ts` and diagnostic tests.
14
+ - `packages/execution/src/types.ts`: optional array of actual command receipts.
15
+ - Runner and CLI adapter receipt propagation, coordinated with WO-48.
16
+
17
+ ## Acceptance
18
+
19
+ - [x] String/null/object/invalid/empty steps and invalid path/fix arguments are rejected before execution, with actionable errors.
20
+ - [x] Valid subsets execute exactly; explicit unavailable checks and zero detected checks cannot report success.
21
+ - [x] No implicit dependency download or empty-test override; actual exit failures remain visible regardless of stdout.
22
+ - [x] Event-loop responsiveness, bounded timeout, Stop, descendant cleanup and sibling isolation are exercised with local fixtures.
23
+ - [x] Actual per-step receipts survive adapter, runner evidence recording and terminal reporting, including mixed outcomes.
24
+ - [x] Record full package verification, scoped commit and push.
25
+
26
+ ## Implemented behavior and focused verification
27
+
28
+ `DiagnosticTool.validateInput` and direct `execute` share the same value-level guard. The production CLI regression reproduces the field's JSON-string `steps` through the real adapter and runner: validation rejects it before the tool body or a process executes. Unsupported or unavailable explicit checks reject the whole request before dispatch. Default detection requires at least one available command.
29
+
30
+ Configured package scripts run through `npm run <step>` with argument arrays; fallback runners must already exist in a local `node_modules/.bin` directory. The tool no longer launches `npx` or adds Jest's `--passWithNoTests`. Local ESLint and Biome receive their respective `--fix` and `--write` flags. Commands execute using `runProcessBuffer` with a copied environment, bounded capture, a per-command timeout, and private process-group cancellation on Unix. Stop drains the owning invocation and prevents later steps; another tool instance remains independent. Build, lint fixes and package scripts retain external-effect authority.
31
+
32
+ `ToolResult.executionReceipts` carries one actual command receipt per attempted step. Completed processes retain their actual exit status. Timeout, cancellation and incomplete capture use null effective/primary status, preserving any raw zero exit in `runnerExitCode`; a cancelled process cannot become verification success merely by handling TERM with `exit(0)`. Earlier completed checks remain separately attributable when a later check fails. No aggregate successful command is invented.
33
+
34
+ Focused verification (2026-09-05, synthetic processes and temporary projects only):
35
+
36
+ - `packages/execution/tests/diagnostic.test.ts`: 35 passing tests.
37
+ - Existing `packages/execution/tests/process-async.test.ts`: 1 passing descendant-cancellation test.
38
+ - `packages/cli/tests/diagnostic-boundary.test.ts`: 3 passing production-boundary tests, including exact two-command terminal reports with exit pairs `[0, 0]` and `[0, 7]`.
39
+ - Execution typecheck passed. Logs: `/tmp/omnius-wo47-diagnostic-tests.log`, `/tmp/omnius-wo47-cli-boundary-tests.log`, `/tmp/omnius-wo47-execution-types.log`.
40
+
41
+ These receipts prove the observed process outcomes, not the completeness of a project's test suite. Configured package scripts remain project-owned commands and may themselves perform network access or modify files; the diagnostic tool does not invent a sandbox or infer an exact changed-file list. Unix descendant ownership is covered on the current host. Publication and live validation remain outside this workorder's repair execution.
42
+
43
+ The complete execution package passed 1,820 tests with 3 pre-existing skips across 163 suites (`/tmp/omnius-remedies-execution-full.log`). The production CLI boundary tests and receipt consumer ship with WO48 so that each source commit has its dependencies.
44
+
45
+ ## Repository acceptance and delivery
46
+
47
+ - Source repair: `d4c37489`, delivered to `origin/main`; coordinated changes: `672bef53 (CLI adapter and report integration)`.
48
+ - Clean rebuild of all workspace packages and final orchestrator rebuild passed.
49
+ - Execution: 1,820 passed / 3 existing skips across 163 suites.
50
+ - CLI: 2,575 passed across 274 suites. Rerun with `--maxWorkers=4 --minWorkers=1` cleared the original five UI import timeouts without changing those tests or their limits.
51
+ - Orchestrator: full run covered 228 suites with 2,847 passed / 1 existing skip / 1 outdated fixture failure. The fixture lacked required `pipefail`; it was corrected, and all 10 tests in that suite then passed. No production source changed after the full run began. This is full-suite coverage plus the corrected suite rerun, not a claim that the initial command exited successfully.
52
+ - Across these packages, all 7,243 distinct non-skipped tests have a passing result on the final implementation. Logs: `/tmp/omnius-remedies-execution-full.log`, `/tmp/omnius-remedies-orchestrator-full.log`, `/tmp/omnius-remedies-receipt-fixture.log`, `/tmp/omnius-remedies-cli-final.log`.
53
+
54
+ Checks used hermetic backend/Telegram transport and synthetic local process/project fixtures. No live inference, service restart, installed-package replacement, Telegram sends or publication were performed. The existing publish staging and unrelated discovery changes were preserved.
@@ -0,0 +1,63 @@
1
+ # WO-48: Retain command evidence in conversation results
2
+
3
+ Status: repository repair complete and delivered to `origin/main`. Publication and live validation remain with the user.
4
+
5
+ ## Finding and root repair
6
+
7
+ Field review F3 found an accepted conversation final with actual check output but an empty structured terminal report. Conversation mode intentionally disables the task ledger; the shared result producer therefore discards evidence before reporting.
8
+
9
+ Maintain report-only evidence for the exact conversation run and task epoch, reusing existing evidence types and immutable artifact storage. Preserve accepted process facts and observations independently of the task ledger. Persist the report snapshot before terminal delivery. Task claims, todos, workboards, mission artifacts and completion requirements must remain governed by their existing authority; enabling reporting cannot grant task or mutation permissions.
10
+
11
+ ## Owned paths
12
+
13
+ - `packages/orchestrator/src/agenticRunner.ts`: evidence producer, run/epoch initialization and terminal commit.
14
+ - `packages/orchestrator/src/conversationReportEvidence.ts` and production-runner/recovery tests.
15
+ - `packages/orchestrator/src/commandEvidenceReceipt.ts`: shared strict receipt parser before nested field access or recovery.
16
+ - `packages/cli/src/tui/tool-adapter.ts`: multiple command receipt preservation.
17
+ - `packages/orchestrator/src/terminalTaskReport.ts`: typed visible limits and successful observation outcomes.
18
+ - Actual mocked Telegram conversation/final-selection tests and adapter tests.
19
+
20
+ ## Acceptance
21
+
22
+ - [x] Actual conversation-mode execution retains typed command and observation evidence in both loops.
23
+ - [x] Actual Telegram chat route delivers a useful report when no separately authored final exists, including failed checks.
24
+ - [x] Task ledger/claims/todos/mission state remain absent; reporting does not change completion authority or tool permissions.
25
+ - [x] New runs, epoch replacement, stale callbacks and scope mismatch cannot inherit earlier report evidence.
26
+ - [x] Immutable persistence/recovery retains exact scoped evidence or reports an explicit gap; no fabricated restored facts.
27
+ - [x] Multi-step diagnostic outcomes are recorded individually, preserving earlier passes and later failure/cancellation.
28
+ - [x] Record full package verification, scoped commit and push.
29
+
30
+ ## Integration design and additional defects found
31
+
32
+ Conversation reporting retains successful reads/searches and typed command outcomes without creating the task completion ledger. Report facts never enter task admission, claim reconciliation or completion readiness. An assessment can complete while truthfully reporting a failed project check. Run and epoch initialization clear earlier evidence, and each tool dispatch captures its report identity before any await so late callbacks cannot populate a replacement report.
33
+
34
+ Multi-command results retain separate process outcomes. The aggregate keeps its own mutation/delivery success; an earlier passing child command cannot make a failed aggregate delivery or mutation successful. The CLI adapter carries every receipt through direct and streaming execution.
35
+
36
+ Actual Telegram regressions exposed a second presentation path: verification-audit gap prose included the tool's raw stdout summary. The terminal report now derives visible gap details from typed commands, paths and stable gap categories. Raw evidence remains inspectable. Successful negative searches and normalized read-preview SIGPIPE outcomes are observations with their original exits retained, rather than invented failures or verification passes.
37
+
38
+ Recovery storage uses immutable linked chunks plus a scoped selector, rather than rewriting the entire evidence history after every tool. Only new evidence is stored, avoiding quadratic use of the shared artifact-store quota on long runs. Recovery validates scope, ordering, identities, receipt shape and integrity before reconstructing report-only facts; it cannot deserialize task authority. Missing or rejected evidence becomes an explicit report limit. Final delivery requires the current evidence to be durably persisted.
39
+
40
+ ## Focused verification
41
+
42
+ - 16 production-runner cases passed, covering both loops; mixed pass/failure/timeout; no task authority; new-run isolation; stale callbacks after epoch replacement; read/search observations; separate aggregate effects; malformed receipt fields with valid siblings; repeated identical attempts; and exact interrupted recovery or an explicit gap without replaying checks. Log: `/tmp/omnius-wo48-runner.log`.
43
+ - 49 helper cases passed, including 100 receipts under a 256 KiB storage budget, immutable old-head recovery, malformed scope/receipt/task-authority rejection, corruption and missing ancestors, count/byte/cycle limits, and publication fault points. The three recovery-limit cases additionally prove that rejected count/byte bounds prevent CAS body reads. Logs: `/tmp/omnius-wo48-report-evidence-helper-tests.log`, `/tmp/omnius-wo48-report-recovery-limits-tests.log`.
44
+ - 36 terminal report tests passed. Log: `/tmp/omnius-wo48-terminal-report.log`.
45
+ - Six actual Telegram conversation cases, three adapter paths and three actual diagnostic-to-runner cases passed in the complete CLI suite. Final CLI: 2,575 passing tests across 274 suites with four workers; original unlimited concurrency hit five existing UI import timeouts. Log: `/tmp/omnius-remedies-cli-final.log`.
46
+ - Clean all-workspace build and final orchestrator rebuild passed. Logs: `/tmp/omnius-remedies-clean-build.log`, `/tmp/omnius-remedies-orchestrator-final-build.log`.
47
+
48
+ The strict receipt parser also exposed an older valid-verification mock missing required `pipefail`; the fixture now supplies the actual contract field, and all ten observed-authority cases pass without weakening readiness assertions (`/tmp/omnius-remedies-receipt-fixture.log`).
49
+
50
+ The review additionally reproduced invalid nested receipt fields throwing before validation, and an incomplete aggregate disappearing behind its passed prerequisite. The recorder now validates before nested field access, preserves unverified evidence and valid siblings, and records the aggregate failure separately. Separate identical array receipts remain distinct executions; only a legacy singleton alias is deduplicated.
51
+
52
+ Recovery is bounded by configured CAS and host recovery limits; it is not unlimited retention or a filesystem-effect transaction. An interruption before a tool observation is published still uses the existing interruption reconciliation. Publication and live validation remain with the user.
53
+
54
+ ## Repository acceptance and delivery
55
+
56
+ - Source repair: `672bef53`, delivered to `origin/main`; coordinated changes: `a3c26eca and d4c37489 (shared command semantics and diagnostic receipts)`.
57
+ - Clean rebuild of all workspace packages and final orchestrator rebuild passed.
58
+ - Execution: 1,820 passed / 3 existing skips across 163 suites.
59
+ - CLI: 2,575 passed across 274 suites. Rerun with `--maxWorkers=4 --minWorkers=1` cleared the original five UI import timeouts without changing those tests or their limits.
60
+ - Orchestrator: full run covered 228 suites with 2,847 passed / 1 existing skip / 1 outdated fixture failure. The fixture lacked required `pipefail`; it was corrected, and all 10 tests in that suite then passed. No production source changed after the full run began. This is full-suite coverage plus the corrected suite rerun, not a claim that the initial command exited successfully.
61
+ - Across these packages, all 7,243 distinct non-skipped tests have a passing result on the final implementation. Logs: `/tmp/omnius-remedies-execution-full.log`, `/tmp/omnius-remedies-orchestrator-full.log`, `/tmp/omnius-remedies-receipt-fixture.log`, `/tmp/omnius-remedies-cli-final.log`.
62
+
63
+ Checks used hermetic backend/Telegram transport and synthetic local process/project fixtures. No live inference, service restart, installed-package replacement, Telegram sends or publication were performed. The existing publish staging and unrelated discovery changes were preserved.
@@ -1,12 +1,12 @@
1
1
  {
2
2
  "name": "omnius",
3
- "version": "1.0.696",
3
+ "version": "1.0.698",
4
4
  "lockfileVersion": 3,
5
5
  "requires": true,
6
6
  "packages": {
7
7
  "": {
8
8
  "name": "omnius",
9
- "version": "1.0.696",
9
+ "version": "1.0.698",
10
10
  "bundleDependencies": [
11
11
  "image-to-ascii"
12
12
  ],
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "omnius",
3
- "version": "1.0.696",
3
+ "version": "1.0.698",
4
4
  "description": "AI coding agent powered by open-source models (Ollama/vLLM) — interactive TUI with agentic tool-calling loop",
5
5
  "type": "module",
6
6
  "main": "./dist/library.js",