omnius 1.0.591 → 1.0.592
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.aiwg/addons/omnius-docs/README.md +15 -1
- package/.aiwg/addons/omnius-docs/manifest.json +28 -68
- package/.aiwg/addons/omnius-docs/skills/agent-failure-recovery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/browser-interaction-validation/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/evidence-directed-delivery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/hardware-evidence-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/omnius-docs/SKILL.md +17 -7
- package/.aiwg/addons/omnius-docs/skills/omnius-inference-docs/SKILL.md +27 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-integration-docs/SKILL.md +21 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-ops-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-realtime-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-sponsor-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-telegram-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-tools-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-version-compatibility-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/runtime-provenance-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/secrets-and-config-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/test-surface-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/workspace-reality-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-rest-docs/README.md +3 -0
- package/.aiwg/addons/omnius-rest-docs/manifest.json +27 -20
- package/.aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md +9 -5
- package/README.md +36 -0
- package/dist/discovery.d.ts +50 -0
- package/dist/index.js +5975 -4021
- package/dist/library.d.ts +7 -0
- package/dist/library.js +950 -0
- package/dist/postinstall-daemon.cjs +18 -0
- package/dist/providerRegistry.d.ts +80 -0
- package/dist/service-version.d.ts +35 -0
- package/docs/.vitepress/config.mts +8 -0
- package/docs/DISCOVERY.json +20224 -0
- package/docs/DISCOVERY.md +648 -0
- package/docs/HANDOFF-crl-encoder-decoder-fix.md +129 -0
- package/docs/agent-memory/INDEX.md +9 -4
- package/docs/agent-memory/index.md +7 -0
- package/docs/concept-relational-language.md +869 -0
- package/docs/context-management-medium-models-proposal.md +449 -0
- package/docs/dedup-false-positive-meta-analysis.md +96 -0
- package/docs/discovery/catalog-overrides.json +724 -0
- package/docs/duplicate-calls-root-cause-analysis.md +91 -0
- package/docs/duplicate-calls-root-cause-deep.md +155 -0
- package/docs/ephemeral-skill-pack-small-context.md +57 -0
- package/docs/explorations/context-window-todo-association.md +156 -0
- package/docs/explorations/todo-association-verify.json +30 -0
- package/docs/explorations/verification-ledger.json +45 -0
- package/docs/explorations/verify-todo-association.sh +30 -0
- package/docs/flowstate.md +806 -0
- package/docs/getting-started/install.md +24 -0
- package/docs/getting-started/model-providers.md +13 -0
- package/docs/guides/agent-integration.md +87 -0
- package/docs/guides/bring-your-own-inference.md +126 -0
- package/docs/guides/tools-and-web-search.md +95 -0
- package/docs/index.md +14 -0
- package/docs/longhaul-35b-workorders.md +496 -0
- package/docs/memory-integration-analysis.md +303 -0
- package/docs/model-capability-awareness-and-multimodal-memory-root-fix.md +799 -0
- package/docs/multimodal-identity-memory-implementation.md +76 -0
- package/docs/omnius-self-edit-eval-2026-06-10.md +169 -0
- package/docs/opencode-agentic-loop-comparison.md +290 -0
- package/docs/operations/security-and-remote-access.md +2 -2
- package/docs/operations/version-compatibility.md +63 -0
- package/docs/proposals/git-progress-tracking-strategy.md +289 -0
- package/docs/proposals/opencode-modules/backendAdapter.ts +443 -0
- package/docs/proposals/opencode-modules/childSession.ts +288 -0
- package/docs/proposals/opencode-modules/compactionAgent.ts +101 -0
- package/docs/proposals/opencode-modules/orchestrator.ts +387 -0
- package/docs/proposals/opencode-modules/runner.ts +258 -0
- package/docs/reference/auth-map.md +87 -196
- package/docs/reference/configuration.md +27 -0
- package/docs/reference/rest-api.md +7 -0
- package/docs/reference/slash-commands.md +125 -2
- package/docs/research/_archived/README.md +18 -0
- package/docs/research/_archived/context_window_attention_model.py +418 -0
- package/docs/research/_archived/context_window_attention_spec.md +55 -0
- package/docs/research/_archived/context_window_attention_weights.json +68 -0
- package/docs/research/k-splanifolds.pdf +0 -0
- package/docs/research/personality-verbosity-control.md +293 -0
- package/docs/rest/INDEX.md +7 -0
- package/docs/rest/QUICKREF.md +18 -0
- package/docs/rest/REST-DOCS-MANIFEST.json +1 -0
- package/docs/rest/auth-and-scopes.md +7 -1
- package/docs/rest/endpoints/discovery.md +44 -0
- package/docs/rest/endpoints/events.md +5 -0
- package/docs/rest/endpoints/tools.md +9 -0
- package/docs/reviews/adversary-system-review.md +42 -0
- package/docs/sana-and-video-generation-integration-plan.md +712 -0
- package/docs/session-diary-llm-training-analysis.md +218 -0
- package/docs/telegram-dmn-curiosity-outreach-scaffold.md +91 -0
- package/docs/telegram-mid-horizon-download-loop-handoff.md +468 -0
- package/docs/telegram-reflection-corpus-integration-plan.md +306 -0
- package/docs/telegram-unified-tooling-architecture.md +332 -0
- package/docs/threat-model.md +868 -0
- package/docs/trajectory-grounding.md +160 -0
- package/docs/voice-flow-architecture.md +489 -0
- package/docs/work-orders/WO-AM-GAPS.md +638 -0
- package/docs/work-orders/daemon-hud-ui-overhaul.md +82 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/INDEX.md +21 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/WORKORDER.md +225 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/WORKORDER.md +198 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/WORKORDER.md +172 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/WORKORDER.md +169 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/WORKORDER.md +189 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/WORKORDER.md +199 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/WORKORDER.md +174 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/WORKORDER.md +226 -0
- package/docs/work-orders/hermes-architecture-deltas/INDEX.md +38 -0
- package/docs/work-orders/omnius-context-engineering-behavior-fixes.md +281 -0
- package/docs/work-orders/telegram-dropbear-context-rca-workorder.md +202 -0
- package/docs/work-orders/world-class-memory-compiler/README.md +162 -0
- package/docs/work-orders/world-class-memory-compiler/TRACKER.md +179 -0
- package/docs/work-orders/world-class-memory-compiler/WO-01-exact-request-budget.md +79 -0
- package/docs/work-orders/world-class-memory-compiler/WO-02-typed-memory-fabric.md +65 -0
- package/docs/work-orders/world-class-memory-compiler/WO-03-dependency-working-set.md +55 -0
- package/docs/work-orders/world-class-memory-compiler/WO-04-inference-memory-compiler.md +67 -0
- package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md +72 -0
- package/docs/work-orders/world-class-memory-compiler/WO-06-temporal-hybrid-retrieval.md +49 -0
- package/docs/work-orders/world-class-memory-compiler/WO-07-evaluation-harness.md +45 -0
- package/docs/work-orders/world-class-memory-compiler/WO-08-rollout-legacy-removal.md +45 -0
- package/docs/x402-remote-inference-plan.md +323 -0
- package/npm-shrinkwrap.json +108 -117
- package/package.json +7 -6
- package/templates/AGENTS.md +6 -0
- package/templates/OMNIUS.md +20 -0
package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md
ADDED
|
@@ -0,0 +1,72 @@
|
|
|
1
|
+
# WO-05 — Artifact Fidelity and Explicit Materialization
|
|
2
|
+
|
|
3
|
+
**Status:** in progress
|
|
4
|
+
**Primary modules:** `packages/execution/src/tools/file-read.ts`,
|
|
5
|
+
`packages/orchestrator/src/evidenceBranch.ts`, `evidenceLedger.ts`,
|
|
6
|
+
`artifactContract.ts`
|
|
7
|
+
**Depends on:** WO-02, WO-03
|
|
8
|
+
|
|
9
|
+
## Problem
|
|
10
|
+
|
|
11
|
+
Mandatory curated branch-extract wrappers made the model distrust first-class
|
|
12
|
+
reads and fall back to shell. Automatic rehydration then injected partial or
|
|
13
|
+
duplicated material without explaining why it was present.
|
|
14
|
+
|
|
15
|
+
## Contract
|
|
16
|
+
|
|
17
|
+
`full_read` is canonical artifact evidence. `branch_extract` is a derived
|
|
18
|
+
artifact linked to source hash/range, requirement coverage, unresolved
|
|
19
|
+
requirements, extractor identity, and confidence. Materialization is an
|
|
20
|
+
explicit graph decision, never an opaque prompt wrapper.
|
|
21
|
+
|
|
22
|
+
## Todos
|
|
23
|
+
|
|
24
|
+
- [x] Persist full-read artifacts with hash, revision, range, fidelity, and
|
|
25
|
+
availability state.
|
|
26
|
+
- [x] Make branch extraction optional (automatic only when the request budget
|
|
27
|
+
cannot safely carry the full body) and callable by the model.
|
|
28
|
+
- [x] Require each extract to account for every requested requirement as
|
|
29
|
+
satisfied with an anchor or unresolved with a bounded reason/recovery.
|
|
30
|
+
- [x] Never mark a partial or capped tool result `not_found`; record truncated
|
|
31
|
+
provenance and deterministic narrowing advice.
|
|
32
|
+
- [ ] Retire the remaining legacy cache-rehydration presentation path after
|
|
33
|
+
materialization telemetry has passed shadow/canary comparison.
|
|
34
|
+
- [x] On content-hash change, supersede old artifact evidence and allow the
|
|
35
|
+
fresh read without stale-read penalties.
|
|
36
|
+
|
|
37
|
+
## Acceptance tests
|
|
38
|
+
|
|
39
|
+
- Twenty oversized files retain source hashes and explicit extract coverage;
|
|
40
|
+
none can return a one-point “complete” extract with open requirements.
|
|
41
|
+
- An edit depending on a read gets the exact body or a visible request for it,
|
|
42
|
+
never a deceptive wrapper.
|
|
43
|
+
- A capped 1 MB grep result is partial success with narrowing, not discovery
|
|
44
|
+
failure.
|
|
45
|
+
- Hash-changing a file permits and records a justified reread.
|
|
46
|
+
|
|
47
|
+
## Live-harness safety gate
|
|
48
|
+
|
|
49
|
+
`live-branch-extract-harness.mjs` creates twenty individual oversized source
|
|
50
|
+
files and performs isolated extraction only. It requires
|
|
51
|
+
`HARNESS_APPROVED_GPU_UUID`; before its first generation it independently
|
|
52
|
+
checks `nvidia-smi`, the selected UUID, and the Ollama service's explicit
|
|
53
|
+
`CUDA_VISIBLE_DEVICES`. All visible service devices must be approved
|
|
54
|
+
A100-class accelerators. A missing, mismatched, or low-capability device is a
|
|
55
|
+
hard failure, never a fallback to CPU or another GPU.
|
|
56
|
+
|
|
57
|
+
The harness now has two explicit requirements per artifact—exact `baudRate`
|
|
58
|
+
and exact `channel`—and fails if either is unresolved, unanchored, or omitted
|
|
59
|
+
from the derived evidence. `HARNESS_REPORT_FILE` writes a structured,
|
|
60
|
+
non-source-bearing checkpoint after each artifact; a terminal report is
|
|
61
|
+
required for acceptance.
|
|
62
|
+
|
|
63
|
+
**Recorded acceptance:** 2026-07-13, `robit/ornith:35b` through Ollama on the
|
|
64
|
+
operator-approved A100 UUID completed 20/20 artifacts. Every artifact had two
|
|
65
|
+
satisfied anchored requirements and no unresolved requirement; the largest
|
|
66
|
+
isolated prompt was 6,080 characters while the parent received at most 1,611
|
|
67
|
+
characters of evidence. The local report is intentionally text-free.
|
|
68
|
+
|
|
69
|
+
## Definition of done
|
|
70
|
+
|
|
71
|
+
Tool tests plus an inference-driven extraction harness on approved A100-class
|
|
72
|
+
hardware demonstrate coverage and source-to-action fidelity.
|
|
@@ -0,0 +1,49 @@
|
|
|
1
|
+
# WO-06 — Temporal Hybrid Retrieval
|
|
2
|
+
|
|
3
|
+
**Status:** in progress
|
|
4
|
+
**Primary modules:** `packages/memory/src/hybridMemoryRetriever.ts`,
|
|
5
|
+
`pprRetrieval.ts`, `temporalGraph.ts`, `memoryGraph.ts`
|
|
6
|
+
**Depends on:** WO-02, WO-03, WO-05
|
|
7
|
+
|
|
8
|
+
## Problem
|
|
9
|
+
|
|
10
|
+
Vector-only retrieval misses exact code locations and versions; lexical-only
|
|
11
|
+
retrieval misses associations; timeless retrieval resurrects obsolete evidence.
|
|
12
|
+
|
|
13
|
+
## Contract
|
|
14
|
+
|
|
15
|
+
Retrieve from a union of exact path/hash/range lookup, lexical search, semantic
|
|
16
|
+
search, graph traversal, and temporal validity filtering. Rerank only records
|
|
17
|
+
allowed by authority and task epoch. Return evidence IDs and abstain where the
|
|
18
|
+
support threshold is not met.
|
|
19
|
+
|
|
20
|
+
## Todos
|
|
21
|
+
|
|
22
|
+
- [x] Add a retrieval query model containing task epoch, active action, claim
|
|
23
|
+
IDs, artifact revision constraints, authority boundary, and token budget.
|
|
24
|
+
- [x] Implement exact-artifact and lexical candidates before vector/graph
|
|
25
|
+
expansion.
|
|
26
|
+
- [x] Add temporal supersession filtering and contradiction surfacing.
|
|
27
|
+
- [x] Use dependency-graph traversal for multi-hop candidates.
|
|
28
|
+
- [x] Add a deterministic reranker with diversity and body-once constraints.
|
|
29
|
+
- [x] Return `abstain` with a specific missing-evidence request instead of
|
|
30
|
+
generating unsupported orientation.
|
|
31
|
+
|
|
32
|
+
The implemented slice is deliberately bounded in `ContextMemoryLedger`: it
|
|
33
|
+
uses exact path/revision/range constraints, optional host semantic scores, a
|
|
34
|
+
six-hop/128-node active-work graph walk, a 32-record/48k-character materialized
|
|
35
|
+
body ceiling, and an explicit `graphTruncated` signal. It is not yet the full
|
|
36
|
+
cross-store/vector retriever or replay-trace quality evaluation required by the
|
|
37
|
+
definition of done.
|
|
38
|
+
|
|
39
|
+
## Acceptance tests
|
|
40
|
+
|
|
41
|
+
- A changed file revision outranks and supersedes its older body.
|
|
42
|
+
- Exact cited ranges outrank semantically similar but unrelated code.
|
|
43
|
+
- A two-hop task-to-claim-to-evidence query returns all required support.
|
|
44
|
+
- Missing support returns abstention rather than fabricated facts.
|
|
45
|
+
|
|
46
|
+
## Definition of done
|
|
47
|
+
|
|
48
|
+
Retrieval tests pass on synthetic temporal graphs and on redacted Omnius trace
|
|
49
|
+
fixtures with measured next-action evidence recall and precision.
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# WO-07 — Replay Evaluation and Adversarial Harness
|
|
2
|
+
|
|
3
|
+
**Status:** in progress
|
|
4
|
+
**Primary modules:** `packages/orchestrator/scripts/`, test fixtures, context
|
|
5
|
+
audit/log serializers
|
|
6
|
+
**Depends on:** WO-01 through WO-06
|
|
7
|
+
|
|
8
|
+
## Problem
|
|
9
|
+
|
|
10
|
+
Compaction quality cannot be inferred from summary fluency or a few unit tests.
|
|
11
|
+
It must be measured against actual long-haul traces, including the MyActuator
|
|
12
|
+
failure modes.
|
|
13
|
+
|
|
14
|
+
## Todos
|
|
15
|
+
|
|
16
|
+
- [ ] Define a redacted, deterministic trace-fixture format containing events,
|
|
17
|
+
artifacts, source revisions, request budgets, expected evidence, and outcomes.
|
|
18
|
+
- [ ] Import representative historical traces without retaining private source
|
|
19
|
+
bodies in committed fixtures.
|
|
20
|
+
- [ ] Implement baseline runners: current behavior, no compaction, pure
|
|
21
|
+
summary, selector-only, and memory compiler.
|
|
22
|
+
- [ ] Measure evidence recall/precision, stale-evidence use, duplicate reads,
|
|
23
|
+
false-not-found, repeated action signatures, completion accuracy, final
|
|
24
|
+
request headroom, latency, cost, and analyst cache hit rate.
|
|
25
|
+
- [ ] Add required adversarial scenarios: 20+ oversized files; partial extract;
|
|
26
|
+
mid-action no-mutation steering; source revision; repeated verifier after
|
|
27
|
+
state change; prompt-like tool output; recap restoration; long exploration.
|
|
28
|
+
- [ ] Exercise small, medium, and large model tiers. Live model harnesses must
|
|
29
|
+
use the approved GPU policy and record selected hardware; no non-capable GPU
|
|
30
|
+
is ever eligible.
|
|
31
|
+
|
|
32
|
+
## Acceptance tests
|
|
33
|
+
|
|
34
|
+
- Harness result is deterministic for a fixed mocked backend and fixture.
|
|
35
|
+
- Every scenario has a failure assertion, not just a happy-path score.
|
|
36
|
+
- A result records the exact configuration, request fingerprint, fixture hash,
|
|
37
|
+
model tier, and compiler version.
|
|
38
|
+
- The MyActuator fixture proves no old-plan mutation follows no-mutation
|
|
39
|
+
steering and no exploration/tool gate is introduced.
|
|
40
|
+
|
|
41
|
+
## Definition of done
|
|
42
|
+
|
|
43
|
+
A committed local harness produces a machine-readable comparison report, and
|
|
44
|
+
the quality gate has objective promotion thresholds rather than a subjective
|
|
45
|
+
“looks better” judgment.
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# WO-08 — Shadow Rollout, Migration, and Legacy Removal
|
|
2
|
+
|
|
3
|
+
**Status:** in progress
|
|
4
|
+
**Primary modules:** orchestrator configuration, telemetry, TUI audit display,
|
|
5
|
+
legacy compaction/rehydration paths
|
|
6
|
+
**Depends on:** WO-07
|
|
7
|
+
|
|
8
|
+
## Rollout contract
|
|
9
|
+
|
|
10
|
+
The strict v2 path is the default request compiler: at the exact 40%-free
|
|
11
|
+
headroom boundary it uses only a validated inference plan, and a failure holds
|
|
12
|
+
the unmodified request. `OMNIUS_MEMORY_COMPILER_MODE=shadow` is an explicit
|
|
13
|
+
rollback/diagnostic mode while comparison and canary evidence is gathered.
|
|
14
|
+
`hold` must never block a read, edit, exploration, or verifier tool call.
|
|
15
|
+
|
|
16
|
+
## Todos
|
|
17
|
+
|
|
18
|
+
- [ ] Add feature flags for ledger dual-write, graph materialization, shadow
|
|
19
|
+
compiler, active compiler, and second opinion; document defaults.
|
|
20
|
+
- [x] Emit body-free TUI/log lifecycle receipts: exact request budget →
|
|
21
|
+
inference dispositions/locators → applied/held/rejected request fingerprint.
|
|
22
|
+
- [ ] Define promotion thresholds from WO-07 and reject rollout if authority,
|
|
23
|
+
fidelity, or tool-freedom regress.
|
|
24
|
+
- [ ] Canary by session/model tier with automatic rollback to untouched history
|
|
25
|
+
on schema/validation/runtime failure.
|
|
26
|
+
- [ ] Run and document a rollback drill.
|
|
27
|
+
- [ ] Remove synthetic recap, silent recovery, heuristic fallback, and stale
|
|
28
|
+
controller injection code only after canary acceptance.
|
|
29
|
+
- [ ] Remove obsolete flags, tests, and docs in the same release; do not leave
|
|
30
|
+
dead legacy behavior silently enabled by environment variables.
|
|
31
|
+
|
|
32
|
+
## Acceptance tests
|
|
33
|
+
|
|
34
|
+
- Shadow mode changes no outgoing request while producing an auditable delta.
|
|
35
|
+
- A malformed delta rolls back before request send and does not affect tool
|
|
36
|
+
availability.
|
|
37
|
+
- Rollback restores the last known valid materialized working set.
|
|
38
|
+
- After legacy deletion, a source scan proves no synthetic recap or automatic
|
|
39
|
+
recovery injection remains on the normal request path.
|
|
40
|
+
|
|
41
|
+
## Definition of done
|
|
42
|
+
|
|
43
|
+
The canary report meets published thresholds, rollback has been exercised, the
|
|
44
|
+
legacy source is removed, and the package build plus focused regression suite
|
|
45
|
+
pass. Deployment remains a separate explicitly approved operation.
|
|
@@ -0,0 +1,323 @@
|
|
|
1
|
+
# x402 Remote Inference Integration — Security Audit & Comprehensive Plan
|
|
2
|
+
|
|
3
|
+
## Part 1: Security Audit of x402 Payment Rails
|
|
4
|
+
|
|
5
|
+
### Summary: 4 CRITICAL, 6 HIGH findings
|
|
6
|
+
|
|
7
|
+
#### CRITICAL
|
|
8
|
+
|
|
9
|
+
**C1: Private key not truly zeroed (nexus.ts L1455, L1901)**
|
|
10
|
+
`privKeyHex = "0".repeat(64)` only rebinds the JS variable — the original string lives on the heap until GC. JavaScript strings are immutable.
|
|
11
|
+
- **Remediation**: Use `Buffer` throughout (not strings). `buffer.fill(0)` overwrites bytes in place. Pass Buffer directly to crypto APIs.
|
|
12
|
+
|
|
13
|
+
**C2: x402-wallet.key is a permanent plaintext key file (nexus.ts L1449-1452)**
|
|
14
|
+
Lives in `.omnius/nexus/` for daemon lifetime. Any backup, snapshot, or directory listing tool exfiltrates the key. `wallet.enc` is meaningless as protection while this file exists.
|
|
15
|
+
- **Remediation**: Pass key to daemon via environment variable or anonymous pipe at spawn time. Delete file after daemon reads it (or never write it to disk).
|
|
16
|
+
|
|
17
|
+
**C3: Budget denylist bypass in doSpend (nexus.ts L1795)**
|
|
18
|
+
`checkBudget(amountSmallest, "transfer:direct", "")` passes empty string for peerId, so `deniedPeers` list is never checked. A blocked peer address can still receive a signed transfer.
|
|
19
|
+
- **Remediation**: Pass `targetAddress` as peerId to `checkBudget()`.
|
|
20
|
+
|
|
21
|
+
**C4: TOCTOU on file permissions (nexus.ts L1451-1452, L1464-1465)**
|
|
22
|
+
`writeFile()` then `chmod()` creates a window where the file is world-readable (default umask).
|
|
23
|
+
- **Remediation**: Use `fs.open(path, 'wx', 0o600)` + `fs.write()` to atomically create with correct permissions.
|
|
24
|
+
|
|
25
|
+
#### HIGH
|
|
26
|
+
|
|
27
|
+
**H1: Scrypt passphrase is predictable (nexus.ts L1440-1443)**
|
|
28
|
+
`hostname():username():nexus-wallet` — anyone with shell access to the machine can derive the key.
|
|
29
|
+
- **Remediation**: Add a user-provided PIN or use OS keyring (libsecret/keychain) when available.
|
|
30
|
+
|
|
31
|
+
**H2: npm version injection in daemon auto-install (nexus.ts ~L1107)**
|
|
32
|
+
`npm view open-agents-nexus version` output interpolated into `execSync`. DNS-hijacked registry response could inject shell commands.
|
|
33
|
+
- **Remediation**: Validate version string against `/^\d+\.\d+\.\d+$/` before interpolation.
|
|
34
|
+
|
|
35
|
+
**H3: No signature verification on spend proof (nexus.ts doSpend)**
|
|
36
|
+
The signed proof in `pending-transfer.json` is not verified before saving. If the signing fails silently, an invalid proof is written to disk and ledger.
|
|
37
|
+
- **Remediation**: Verify signature with `verifyTypedData` before writing proof or ledger entry.
|
|
38
|
+
|
|
39
|
+
**H4: Daemon x402 config accepts arbitrary ALCHEMY_API_KEY from env**
|
|
40
|
+
The daemon script reads `process.env.ALCHEMY_API_KEY` and passes it to the NexusClient x402 config. If the daemon runs in a shared environment, this could be exfiltrated.
|
|
41
|
+
- **Remediation**: Validate API key format before passing; consider injecting only via the spawn environment, not inherited env.
|
|
42
|
+
|
|
43
|
+
**H5: EIP-3009 nonce is random but not stored for replay detection**
|
|
44
|
+
`doSpend` generates a random nonce for each transfer but doesn't track used nonces. If a malicious peer resubmits a proof before the original submission, the user could see unexpected behavior.
|
|
45
|
+
- **Remediation**: USDC contract itself prevents nonce replay on-chain. Low risk in practice but should be documented.
|
|
46
|
+
|
|
47
|
+
**H6: No rate limiting on spend action**
|
|
48
|
+
An LLM can be prompt-injected to call `spend` in a loop. Budget policy catches per-day limits but the circuit breaker requires an RPC call that could timeout.
|
|
49
|
+
- **Remediation**: Add a local cooldown (e.g., minimum 5s between spend calls).
|
|
50
|
+
|
|
51
|
+
#### MEDIUM / LOW
|
|
52
|
+
|
|
53
|
+
- **M1**: `containsKeyMaterial` regex doesn't catch Base58 private keys or mnemonic phrases
|
|
54
|
+
- **M2**: Ledger entries written without signing — anyone with file access can forge entries
|
|
55
|
+
- **M3**: Budget policy file (`budget.json`) is not integrity-protected
|
|
56
|
+
- **L1**: No audit log of budget policy changes
|
|
57
|
+
- **L2**: Daemon log may contain sensitive peer IDs (useful for correlation attacks)
|
|
58
|
+
|
|
59
|
+
### Recommended Priority
|
|
60
|
+
|
|
61
|
+
1. ~~**Fix C3 immediately**~~ DONE — `doSpend` now passes `targetAddress` to `checkBudget()`
|
|
62
|
+
2. ~~**Fix C4**~~ DONE — all 3 key/wallet file writes use `fsOpen(path, "w", 0o600)` for atomic creation
|
|
63
|
+
3. **Plan C2 remediation** (daemon key delivery — architectural change, defer to next sprint)
|
|
64
|
+
4. **Document C1** (JS GC limitation — no perfect fix, but can improve with Buffer)
|
|
65
|
+
|
|
66
|
+
---
|
|
67
|
+
|
|
68
|
+
## Part 2: Current State Assessment
|
|
69
|
+
|
|
70
|
+
### What Exists Today
|
|
71
|
+
|
|
72
|
+
| Layer | Component | Status | Remote-Ready? |
|
|
73
|
+
|-------|-----------|--------|---------------|
|
|
74
|
+
| **Agent Loop** | `AgenticRunner` | Production | No — hardcoded local backend |
|
|
75
|
+
| **Backend Interface** | `AgenticBackend` | Production | Yes — clean interface, pluggable |
|
|
76
|
+
| **Backend Impl** | `OllamaAgenticBackend` | Production | Local only |
|
|
77
|
+
| **P2P Transport** | `NexusTool` (daemon-based) | Production | Yes — invoke_capability works |
|
|
78
|
+
| **P2P Mesh** | `PeerMesh` (WebSocket) | Built, not wired | Yes — gossip, heartbeat, capabilities |
|
|
79
|
+
| **Inference Router** | `InferenceRouter` | Built, not wired | Yes — trust scoring, secret redaction |
|
|
80
|
+
| **Secret Vault** | `SecretVault` | Built, not wired | Yes — OMNIUS_VAR placeholder system |
|
|
81
|
+
| **x402 Payments** | Wallet + spend + ledger | Production | Yes — EIP-3009, budget policy |
|
|
82
|
+
| **Sub-Agent** | `OpenCodeTool` | Production | No — spawns local subprocess |
|
|
83
|
+
| **Call Sub-Agent** | `CallSubAgent` | Production | No — uses parent's backend |
|
|
84
|
+
|
|
85
|
+
### Key Architectural Facts
|
|
86
|
+
|
|
87
|
+
1. **`AgenticBackend` is the integration point** — any implementation that satisfies `chatCompletion()` can drive the agent loop
|
|
88
|
+
2. **The InferenceRouter already handles tool-calling** — `P2PInferRequest` includes `tools` array, `P2PInferResponse` includes `toolCalls`
|
|
89
|
+
3. **Secret redaction is automatic** — vault scans all text for known values, replaces with `{{OMNIUS_VAR_*}}`, injects back on response
|
|
90
|
+
4. **Trust tiers control redaction depth** — LOCAL (no redaction), TEE (minimal), VERIFIED (standard), PUBLIC (full)
|
|
91
|
+
5. **The gap is "last mile" wiring** — InferenceRouter exists but is never called from AgenticRunner
|
|
92
|
+
|
|
93
|
+
---
|
|
94
|
+
|
|
95
|
+
## Part 3: The Four Scenarios Evaluated
|
|
96
|
+
|
|
97
|
+
### Scenario 1: Entire Stack Defers to Remote Inference
|
|
98
|
+
*"The whole agent runs on someone else's GPU"*
|
|
99
|
+
|
|
100
|
+
**How it would work**: Replace `OllamaAgenticBackend` with `NexusAgenticBackend` at CLI startup. Every `chatCompletion()` call routes through InferenceRouter to a remote peer.
|
|
101
|
+
|
|
102
|
+
**Already addressed?** Partially. The `AgenticBackend` interface supports this. InferenceRouter handles tool-calling. SecretVault protects secrets. What's missing is:
|
|
103
|
+
- `NexusAgenticBackend` class that adapts InferenceRouter → AgenticBackend interface
|
|
104
|
+
- CLI flag: `--backend nexus` or `--remote-peer 12D3KooW...`
|
|
105
|
+
- x402 budget integration (each chatCompletion costs money)
|
|
106
|
+
|
|
107
|
+
**Attractiveness**: Medium. Useful for headless agents on Raspberry Pi / VPS with no GPU. But latency and trust concerns make it less appealing for primary development.
|
|
108
|
+
|
|
109
|
+
**Security**: SecretVault handles credential safety. Trust tiers control exposure. x402 budget prevents runaway spend. This is actually the **safest** remote scenario.
|
|
110
|
+
|
|
111
|
+
### Scenario 2: Specific Tasks Routed to Remote Models Transiently
|
|
112
|
+
*"I need a 70B model for this one hard coding problem, then back to local 27B"*
|
|
113
|
+
|
|
114
|
+
**How it would work**: AgenticRunner's tool loop detects a "hard" task (or user explicitly requests), temporarily routes to a remote peer with the needed model, gets the response, returns to local inference.
|
|
115
|
+
|
|
116
|
+
**Already addressed?** No. The backend is currently immutable during a task. But the infrastructure is ready:
|
|
117
|
+
- InferenceRouter.`infer(model, messages)` is a one-shot call
|
|
118
|
+
- Could be wrapped as a tool: `nexus(action='remote_infer', model='llama3.3:70b', prompt='...')`
|
|
119
|
+
- Or implemented as backend fallback: local → timeout/quality check → remote
|
|
120
|
+
|
|
121
|
+
**Attractiveness**: HIGH. This is the "superpower" use case. A $200 laptop running 8B can seamlessly tap into a 122B model on the mesh for complex tasks, paying $0.001 per request.
|
|
122
|
+
|
|
123
|
+
**Security**: Medium risk. The "hard task" might contain the most sensitive context. SecretVault mitigates this. Budget policy caps per-invoke spend.
|
|
124
|
+
|
|
125
|
+
### Scenario 3: Sub-Agents Delegated to Remote Inference
|
|
126
|
+
*"Spawn a sub-agent that runs entirely on a remote peer's GPU"*
|
|
127
|
+
|
|
128
|
+
**How it would work**: When spawning a sub-agent, specify a remote backend:
|
|
129
|
+
```
|
|
130
|
+
sub_agent(task='Review this PR', backend='nexus', model='qwen3.5:122b', peer='12D3KooW...')
|
|
131
|
+
```
|
|
132
|
+
The sub-agent's entire AgenticRunner loop runs against the remote peer.
|
|
133
|
+
|
|
134
|
+
**Already addressed?** Partially. CallSubAgent already creates independent AgenticRunner instances. The pattern exists. What's missing:
|
|
135
|
+
- Sub-agent tool that accepts `backend` parameter
|
|
136
|
+
- NexusAgenticBackend (same as Scenario 1)
|
|
137
|
+
- Result return across the network boundary
|
|
138
|
+
- x402 payment for multi-turn conversations (not just single invoke)
|
|
139
|
+
|
|
140
|
+
**Attractiveness**: VERY HIGH. This is the marketplace killer feature. An agent can "hire" specialized remote agents for specific skills. A coding agent sends a security review sub-task to a peer running a security-specialized model.
|
|
141
|
+
|
|
142
|
+
**Security**: Lower risk than Scenario 2 — sub-agent context is scoped to just the delegated task. SecretVault can enforce stricter redaction for sub-agent contexts.
|
|
143
|
+
|
|
144
|
+
### Scenario 4: Interlaced Remote Inference (Any Point in Chain)
|
|
145
|
+
*"Mid-conversation, seamlessly route any individual LLM call to any peer"*
|
|
146
|
+
|
|
147
|
+
**How it would work**: Backend becomes a router, not a fixed endpoint. Each `chatCompletion()` call evaluates:
|
|
148
|
+
1. Is a local model available and capable? → Use local
|
|
149
|
+
2. Is a remote peer better (larger model, lower latency, specific capability)? → Route via InferenceRouter
|
|
150
|
+
3. Apply budget check before routing
|
|
151
|
+
4. Redact secrets, send, inject on response
|
|
152
|
+
|
|
153
|
+
**Already addressed?** The InferenceRouter's scoring formula already supports this:
|
|
154
|
+
```
|
|
155
|
+
score = trustWeight * (1 / (1 + latency/100)) * (1 - load) * modelMatch
|
|
156
|
+
```
|
|
157
|
+
What's missing: the "hybrid backend" that dynamically chooses local vs remote per-call.
|
|
158
|
+
|
|
159
|
+
**Attractiveness**: HIGHEST. This is the most flexible and the end-state vision. But also the most complex to implement correctly.
|
|
160
|
+
|
|
161
|
+
**Security**: Highest risk — any message in the conversation might be sent remotely. Requires robust SecretVault with comprehensive secret detection. Trust tier enforcement is critical.
|
|
162
|
+
|
|
163
|
+
---
|
|
164
|
+
|
|
165
|
+
## Part 4: Recommended Architecture — "The Mix" (Progressive Implementation)
|
|
166
|
+
|
|
167
|
+
### Phase 1: NexusAgenticBackend (enables Scenarios 1 & 3)
|
|
168
|
+
**Effort**: Medium | **Impact**: High | **Timeline**: This sprint
|
|
169
|
+
|
|
170
|
+
Create `NexusAgenticBackend` that implements `AgenticBackend`:
|
|
171
|
+
|
|
172
|
+
```typescript
|
|
173
|
+
export class NexusAgenticBackend implements AgenticBackend {
|
|
174
|
+
constructor(
|
|
175
|
+
private router: InferenceRouter,
|
|
176
|
+
private model: string,
|
|
177
|
+
private budgetChecker?: (cost: number) => Promise<boolean>,
|
|
178
|
+
) {}
|
|
179
|
+
|
|
180
|
+
async chatCompletion(request: ChatCompletionRequest): Promise<ChatCompletionResponse> {
|
|
181
|
+
// 1. Estimate cost from token count
|
|
182
|
+
// 2. Check budget
|
|
183
|
+
// 3. Route via InferenceRouter (handles redaction + peer selection)
|
|
184
|
+
// 4. Map P2PInferResponse → ChatCompletionResponse
|
|
185
|
+
// 5. Write ledger entry
|
|
186
|
+
}
|
|
187
|
+
}
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
**What this unlocks**:
|
|
191
|
+
- `omnius run --backend nexus --model qwen3.5:122b` → full remote agent
|
|
192
|
+
- Sub-agents with `backend: "nexus"` → delegated remote execution
|
|
193
|
+
- `/p2p start` + `/p2p connect` → mesh is live, inference is routed
|
|
194
|
+
|
|
195
|
+
### Phase 2: Remote Inference Tool (enables Scenario 2)
|
|
196
|
+
**Effort**: Low | **Impact**: Very High | **Timeline**: This sprint
|
|
197
|
+
|
|
198
|
+
Add a `remote_infer` action to the nexus tool:
|
|
199
|
+
|
|
200
|
+
```typescript
|
|
201
|
+
case "remote_infer":
|
|
202
|
+
// 1. Find best peer for requested model
|
|
203
|
+
// 2. Budget check
|
|
204
|
+
// 3. Route single inference call via InferenceRouter
|
|
205
|
+
// 4. Return result to agent
|
|
206
|
+
// Agent stays on local model but can "reach out" for specific questions
|
|
207
|
+
```
|
|
208
|
+
|
|
209
|
+
This is the easiest win — a single nexus action that any agent can call. No backend swapping needed. The agent decides when to use remote inference, just like calling any other tool.
|
|
210
|
+
|
|
211
|
+
**What this unlocks**:
|
|
212
|
+
- Agent running on 8B can call `nexus(action='remote_infer', model='qwen3.5:70b', prompt='Complex analysis...')`
|
|
213
|
+
- Budget-checked per-call
|
|
214
|
+
- Secret-safe via vault
|
|
215
|
+
- Agent retains autonomy — it chooses when to use remote help
|
|
216
|
+
|
|
217
|
+
### Phase 3: Hybrid Backend (enables Scenario 4)
|
|
218
|
+
**Effort**: High | **Impact**: Highest | **Timeline**: Next sprint
|
|
219
|
+
|
|
220
|
+
Create `HybridAgenticBackend` that dynamically routes:
|
|
221
|
+
|
|
222
|
+
```typescript
|
|
223
|
+
export class HybridAgenticBackend implements AgenticBackend {
|
|
224
|
+
constructor(
|
|
225
|
+
private localBackend: OllamaAgenticBackend,
|
|
226
|
+
private remoteRouter: InferenceRouter,
|
|
227
|
+
private policy: RoutingPolicy,
|
|
228
|
+
) {}
|
|
229
|
+
|
|
230
|
+
async chatCompletion(request): Promise<Response> {
|
|
231
|
+
const route = this.policy.decide(request, this.localBackend, this.remoteRouter);
|
|
232
|
+
if (route === 'local') return this.localBackend.chatCompletion(request);
|
|
233
|
+
return this.nexusBackend.chatCompletion(request); // via InferenceRouter
|
|
234
|
+
}
|
|
235
|
+
}
|
|
236
|
+
```
|
|
237
|
+
|
|
238
|
+
**RoutingPolicy** decides based on:
|
|
239
|
+
- Model requirements (request asks for capability local model can't provide)
|
|
240
|
+
- Token count (large context → route to peer with bigger context window)
|
|
241
|
+
- Load (local GPU saturated → overflow to mesh)
|
|
242
|
+
- Cost (local is free, remote costs money → prefer local unless quality difference is high)
|
|
243
|
+
- User preference (explicit `/remote on` toggle)
|
|
244
|
+
|
|
245
|
+
### Phase 4: Marketplace Dynamics
|
|
246
|
+
**Effort**: Medium | **Impact**: Network effect | **Timeline**: Post-MVP
|
|
247
|
+
|
|
248
|
+
1. **Provider Dashboard**: `nexus(action='provider_stats')` — earnings, requests served, uptime
|
|
249
|
+
2. **Reputation System**: Track successful invocations, response quality, latency consistency
|
|
250
|
+
3. **Discovery Registry**: Public capability index so agents can find providers without prior connection
|
|
251
|
+
4. **Tiered Pricing**: Providers set per-model rates; consumers see a unified pricing menu
|
|
252
|
+
5. **SLA Guarantees**: Timeout → automatic failover to next-best peer; refund on failure
|
|
253
|
+
|
|
254
|
+
---
|
|
255
|
+
|
|
256
|
+
## Part 5: What Makes This Marketplace Attractive
|
|
257
|
+
|
|
258
|
+
### For Providers (GPU Owners)
|
|
259
|
+
- **Passive income**: Expose idle GPU capacity, earn USDC while sleeping
|
|
260
|
+
- **Zero config**: `omnius run` → `nexus connect` → `nexus expose --margin 0.3` → earning
|
|
261
|
+
- **x402 automatic payments**: No invoicing, no manual settlement. Payment flows with each request
|
|
262
|
+
- **Trust control**: Choose who can access your models (TEE, verified, public)
|
|
263
|
+
- **Usage metering**: Full audit trail in metering.jsonl + ledger.jsonl
|
|
264
|
+
|
|
265
|
+
### For Consumers (Agent Operators)
|
|
266
|
+
- **Access any model**: Your 8B laptop can tap into 122B models on the mesh
|
|
267
|
+
- **Budget safety**: Daily limits, per-invoke caps, circuit breaker — impossible to overspend
|
|
268
|
+
- **Secret safety**: SecretVault auto-redacts credentials before any request leaves your machine
|
|
269
|
+
- **Seamless**: Agent doesn't know it's using remote inference — same tool-calling loop
|
|
270
|
+
- **Transient or persistent**: Single question or entire sub-agent workflow — your choice
|
|
271
|
+
|
|
272
|
+
### For the Network
|
|
273
|
+
- **Self-reinforcing**: More providers → better model selection → more consumers → more revenue → more providers
|
|
274
|
+
- **Anti-centralization**: No single point of failure or control. Any node can be provider AND consumer
|
|
275
|
+
- **Trust graduated**: Start with `public` trust (full redaction), build to `verified`, eventually `tee`
|
|
276
|
+
- **Economic alignment**: x402 ensures providers are compensated, consumers get value, network grows
|
|
277
|
+
|
|
278
|
+
---
|
|
279
|
+
|
|
280
|
+
## Part 6: Implementation Order
|
|
281
|
+
|
|
282
|
+
| Step | Description | Files | Depends On |
|
|
283
|
+
|------|-------------|-------|------------|
|
|
284
|
+
| ~~**0**~~ | ~~Fix C3 budget bypass~~ DONE + C4 TOCTOU fix | nexus.ts | — |
|
|
285
|
+
| **1a** | Wire InferenceRouter into interactive.ts | interactive.ts | Already built |
|
|
286
|
+
| **1b** | Create NexusAgenticBackend | orchestrator/src/nexusBackend.ts | 1a |
|
|
287
|
+
| **1c** | Add `--backend nexus` to CLI | cli/src/config.ts, run.ts | 1b |
|
|
288
|
+
| ~~**2**~~ | ~~Add `remote_infer` action to nexus tool~~ DONE | nexus.ts | 1a |
|
|
289
|
+
| **3a** | Create HybridAgenticBackend | orchestrator/src/hybridBackend.ts | 1b |
|
|
290
|
+
| **3b** | RoutingPolicy with load/cost/capability logic | orchestrator/src/routingPolicy.ts | 3a |
|
|
291
|
+
| **4** | Sub-agent with backend selection | execution/src/tools/ (new or modified) | 1b |
|
|
292
|
+
| **5** | Provider dashboard + reputation | nexus.ts (new actions) | 2 |
|
|
293
|
+
|
|
294
|
+
### Immediate Next Steps (This Sprint)
|
|
295
|
+
|
|
296
|
+
1. ~~**Fix C3**~~ DONE — `targetAddress` now passed to `checkBudget()` in doSpend
|
|
297
|
+
2. ~~**Fix C4**~~ DONE — atomic file creation with `fsOpen(path, "w", 0o600)` for all key/wallet files
|
|
298
|
+
3. ~~**Step 2**~~ DONE — `remote_infer` action with auto-discovery + explicit peer + budget + ledger + error handling (23/23 eval tests pass)
|
|
299
|
+
4. **Step 1b** — NexusAgenticBackend (unlocks Scenarios 1 & 3)
|
|
300
|
+
5. **Step 1c** — CLI flag for nexus backend
|
|
301
|
+
|
|
302
|
+
### What's Already Built vs What's Needed
|
|
303
|
+
|
|
304
|
+
```
|
|
305
|
+
BUILT (just needs wiring):
|
|
306
|
+
├── InferenceRouter (trust-scored peer selection)
|
|
307
|
+
├── SecretVault (automatic credential protection)
|
|
308
|
+
├── PeerMesh (WebSocket gossip mesh)
|
|
309
|
+
├── x402 payment rails (wallet + spend + ledger + budget)
|
|
310
|
+
├── AgenticBackend interface (clean abstraction)
|
|
311
|
+
└── P2P types (InferRequest/Response with tool-calling)
|
|
312
|
+
|
|
313
|
+
DONE:
|
|
314
|
+
├── remote_infer nexus action (auto-discover + invoke + budget + ledger) ✓
|
|
315
|
+
|
|
316
|
+
NEEDS BUILDING:
|
|
317
|
+
├── NexusAgenticBackend (adapter: InferenceRouter → AgenticBackend)
|
|
318
|
+
├── HybridAgenticBackend (local + remote routing)
|
|
319
|
+
├── RoutingPolicy (cost/capability/load decision engine)
|
|
320
|
+
└── CLI integration (--backend nexus, /p2p infer command)
|
|
321
|
+
```
|
|
322
|
+
|
|
323
|
+
The remarkable thing is that **70% of the infrastructure is already built**. The remaining work is primarily wiring and integration — connecting InferenceRouter to AgenticRunner, and adding the CLI/tool surface for agents to use it.
|