omnius 1.0.591 → 1.0.592

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. package/.aiwg/addons/omnius-docs/README.md +15 -1
  2. package/.aiwg/addons/omnius-docs/manifest.json +28 -68
  3. package/.aiwg/addons/omnius-docs/skills/agent-failure-recovery/SKILL.md +2 -1
  4. package/.aiwg/addons/omnius-docs/skills/browser-interaction-validation/SKILL.md +2 -1
  5. package/.aiwg/addons/omnius-docs/skills/evidence-directed-delivery/SKILL.md +2 -1
  6. package/.aiwg/addons/omnius-docs/skills/hardware-evidence-audit/SKILL.md +2 -1
  7. package/.aiwg/addons/omnius-docs/skills/omnius-docs/SKILL.md +17 -7
  8. package/.aiwg/addons/omnius-docs/skills/omnius-inference-docs/SKILL.md +27 -0
  9. package/.aiwg/addons/omnius-docs/skills/omnius-integration-docs/SKILL.md +21 -0
  10. package/.aiwg/addons/omnius-docs/skills/omnius-ops-docs/SKILL.md +2 -0
  11. package/.aiwg/addons/omnius-docs/skills/omnius-realtime-docs/SKILL.md +2 -0
  12. package/.aiwg/addons/omnius-docs/skills/omnius-sponsor-docs/SKILL.md +2 -0
  13. package/.aiwg/addons/omnius-docs/skills/omnius-telegram-docs/SKILL.md +2 -0
  14. package/.aiwg/addons/omnius-docs/skills/omnius-tools-docs/SKILL.md +23 -0
  15. package/.aiwg/addons/omnius-docs/skills/omnius-version-compatibility-docs/SKILL.md +23 -0
  16. package/.aiwg/addons/omnius-docs/skills/runtime-provenance-audit/SKILL.md +2 -1
  17. package/.aiwg/addons/omnius-docs/skills/secrets-and-config-audit/SKILL.md +2 -1
  18. package/.aiwg/addons/omnius-docs/skills/test-surface-audit/SKILL.md +2 -1
  19. package/.aiwg/addons/omnius-docs/skills/workspace-reality-audit/SKILL.md +2 -1
  20. package/.aiwg/addons/omnius-rest-docs/README.md +3 -0
  21. package/.aiwg/addons/omnius-rest-docs/manifest.json +27 -20
  22. package/.aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md +9 -5
  23. package/README.md +36 -0
  24. package/dist/discovery.d.ts +50 -0
  25. package/dist/index.js +5975 -4021
  26. package/dist/library.d.ts +7 -0
  27. package/dist/library.js +950 -0
  28. package/dist/postinstall-daemon.cjs +18 -0
  29. package/dist/providerRegistry.d.ts +80 -0
  30. package/dist/service-version.d.ts +35 -0
  31. package/docs/.vitepress/config.mts +8 -0
  32. package/docs/DISCOVERY.json +20224 -0
  33. package/docs/DISCOVERY.md +648 -0
  34. package/docs/HANDOFF-crl-encoder-decoder-fix.md +129 -0
  35. package/docs/agent-memory/INDEX.md +9 -4
  36. package/docs/agent-memory/index.md +7 -0
  37. package/docs/concept-relational-language.md +869 -0
  38. package/docs/context-management-medium-models-proposal.md +449 -0
  39. package/docs/dedup-false-positive-meta-analysis.md +96 -0
  40. package/docs/discovery/catalog-overrides.json +724 -0
  41. package/docs/duplicate-calls-root-cause-analysis.md +91 -0
  42. package/docs/duplicate-calls-root-cause-deep.md +155 -0
  43. package/docs/ephemeral-skill-pack-small-context.md +57 -0
  44. package/docs/explorations/context-window-todo-association.md +156 -0
  45. package/docs/explorations/todo-association-verify.json +30 -0
  46. package/docs/explorations/verification-ledger.json +45 -0
  47. package/docs/explorations/verify-todo-association.sh +30 -0
  48. package/docs/flowstate.md +806 -0
  49. package/docs/getting-started/install.md +24 -0
  50. package/docs/getting-started/model-providers.md +13 -0
  51. package/docs/guides/agent-integration.md +87 -0
  52. package/docs/guides/bring-your-own-inference.md +126 -0
  53. package/docs/guides/tools-and-web-search.md +95 -0
  54. package/docs/index.md +14 -0
  55. package/docs/longhaul-35b-workorders.md +496 -0
  56. package/docs/memory-integration-analysis.md +303 -0
  57. package/docs/model-capability-awareness-and-multimodal-memory-root-fix.md +799 -0
  58. package/docs/multimodal-identity-memory-implementation.md +76 -0
  59. package/docs/omnius-self-edit-eval-2026-06-10.md +169 -0
  60. package/docs/opencode-agentic-loop-comparison.md +290 -0
  61. package/docs/operations/security-and-remote-access.md +2 -2
  62. package/docs/operations/version-compatibility.md +63 -0
  63. package/docs/proposals/git-progress-tracking-strategy.md +289 -0
  64. package/docs/proposals/opencode-modules/backendAdapter.ts +443 -0
  65. package/docs/proposals/opencode-modules/childSession.ts +288 -0
  66. package/docs/proposals/opencode-modules/compactionAgent.ts +101 -0
  67. package/docs/proposals/opencode-modules/orchestrator.ts +387 -0
  68. package/docs/proposals/opencode-modules/runner.ts +258 -0
  69. package/docs/reference/auth-map.md +87 -196
  70. package/docs/reference/configuration.md +27 -0
  71. package/docs/reference/rest-api.md +7 -0
  72. package/docs/reference/slash-commands.md +125 -2
  73. package/docs/research/_archived/README.md +18 -0
  74. package/docs/research/_archived/context_window_attention_model.py +418 -0
  75. package/docs/research/_archived/context_window_attention_spec.md +55 -0
  76. package/docs/research/_archived/context_window_attention_weights.json +68 -0
  77. package/docs/research/k-splanifolds.pdf +0 -0
  78. package/docs/research/personality-verbosity-control.md +293 -0
  79. package/docs/rest/INDEX.md +7 -0
  80. package/docs/rest/QUICKREF.md +18 -0
  81. package/docs/rest/REST-DOCS-MANIFEST.json +1 -0
  82. package/docs/rest/auth-and-scopes.md +7 -1
  83. package/docs/rest/endpoints/discovery.md +44 -0
  84. package/docs/rest/endpoints/events.md +5 -0
  85. package/docs/rest/endpoints/tools.md +9 -0
  86. package/docs/reviews/adversary-system-review.md +42 -0
  87. package/docs/sana-and-video-generation-integration-plan.md +712 -0
  88. package/docs/session-diary-llm-training-analysis.md +218 -0
  89. package/docs/telegram-dmn-curiosity-outreach-scaffold.md +91 -0
  90. package/docs/telegram-mid-horizon-download-loop-handoff.md +468 -0
  91. package/docs/telegram-reflection-corpus-integration-plan.md +306 -0
  92. package/docs/telegram-unified-tooling-architecture.md +332 -0
  93. package/docs/threat-model.md +868 -0
  94. package/docs/trajectory-grounding.md +160 -0
  95. package/docs/voice-flow-architecture.md +489 -0
  96. package/docs/work-orders/WO-AM-GAPS.md +638 -0
  97. package/docs/work-orders/daemon-hud-ui-overhaul.md +82 -0
  98. package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/INDEX.md +21 -0
  99. package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/WORKORDER.md +225 -0
  100. package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/INDEX.md +20 -0
  101. package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/WORKORDER.md +198 -0
  102. package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/INDEX.md +19 -0
  103. package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/WORKORDER.md +172 -0
  104. package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/INDEX.md +19 -0
  105. package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/WORKORDER.md +169 -0
  106. package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/INDEX.md +22 -0
  107. package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/WORKORDER.md +189 -0
  108. package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/INDEX.md +22 -0
  109. package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/WORKORDER.md +199 -0
  110. package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/INDEX.md +20 -0
  111. package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/WORKORDER.md +174 -0
  112. package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/INDEX.md +22 -0
  113. package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/WORKORDER.md +226 -0
  114. package/docs/work-orders/hermes-architecture-deltas/INDEX.md +38 -0
  115. package/docs/work-orders/omnius-context-engineering-behavior-fixes.md +281 -0
  116. package/docs/work-orders/telegram-dropbear-context-rca-workorder.md +202 -0
  117. package/docs/work-orders/world-class-memory-compiler/README.md +162 -0
  118. package/docs/work-orders/world-class-memory-compiler/TRACKER.md +179 -0
  119. package/docs/work-orders/world-class-memory-compiler/WO-01-exact-request-budget.md +79 -0
  120. package/docs/work-orders/world-class-memory-compiler/WO-02-typed-memory-fabric.md +65 -0
  121. package/docs/work-orders/world-class-memory-compiler/WO-03-dependency-working-set.md +55 -0
  122. package/docs/work-orders/world-class-memory-compiler/WO-04-inference-memory-compiler.md +67 -0
  123. package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md +72 -0
  124. package/docs/work-orders/world-class-memory-compiler/WO-06-temporal-hybrid-retrieval.md +49 -0
  125. package/docs/work-orders/world-class-memory-compiler/WO-07-evaluation-harness.md +45 -0
  126. package/docs/work-orders/world-class-memory-compiler/WO-08-rollout-legacy-removal.md +45 -0
  127. package/docs/x402-remote-inference-plan.md +323 -0
  128. package/npm-shrinkwrap.json +108 -117
  129. package/package.json +7 -6
  130. package/templates/AGENTS.md +6 -0
  131. package/templates/OMNIUS.md +20 -0
@@ -0,0 +1,293 @@
1
+ # LLM Personality Guidance & Verbosity Control: Research Synthesis
2
+
3
+ **Date**: 2026-03-13
4
+ **Purpose**: Literature review and integration strategy for transient personality/verbosity steering in the Omnius framework
5
+
6
+ ---
7
+
8
+ ## 1. Executive Summary
9
+
10
+ This document synthesizes research on controlling LLM output style — specifically verbosity vs. conciseness — through system prompt engineering, persona/personality steering, and length control mechanisms. The findings are organized for direct integration into the Omnius agentic framework, where the system prompt and context engineering pipeline can transiently influence agent behavior without fine-tuning.
11
+
12
+ **Key insight**: Prompt-level personality steering (natural language instructions) categorically dominates activation-level interventions. Explicit system prompt instructions override all other behavioral signals, making prompt engineering the most practical lever for our framework.
13
+
14
+ ---
15
+
16
+ ## 2. Literature & Provenance
17
+
18
+ ### 2.1 Core Papers
19
+
20
+ | # | Paper | Authors | Venue | Year | Key Finding |
21
+ |---|-------|---------|-------|------|-------------|
22
+ | 1 | [The Prompt Report: A Systematic Survey of Prompting Techniques](https://arxiv.org/abs/2406.06608) | Schulhoff et al. (32 authors) | arXiv | 2024 | Taxonomy of 58 prompting techniques; role/persona prompting in §2.2 |
23
+ | 2 | [Same Task, More Tokens](https://arxiv.org/abs/2402.14848) | Levy, Jacoby & Goldberg | ACL 2024 | 2024 | LLM reasoning degrades ~3,000 tokens; sweet spot 150–300 words |
24
+ | 3 | [Lost in the Middle](https://arxiv.org/abs/2307.03172) | Liu, Lin, Hewitt et al. | TACL 2024 | 2024 | U-shaped attention bias; >30% accuracy drop for mid-context info |
25
+ | 4 | [Linear Personality Probing and Steering in LLMs](https://arxiv.org/html/2512.17639v1) | (Big Five study) | arXiv | 2025 | Linear directions can probe personality but explicit prompts override steering vectors entirely |
26
+ | 5 | [SAC: Style Adjective Continuation](https://arxiv.org/abs/2506.20993) | (16PF framework) | arXiv | 2026 | Continuous 1–5 trait intensity via adjective anchoring; 5 behavioral dimensions |
27
+ | 6 | [Precise Length Control in LLMs (LDPE)](https://arxiv.org/abs/2412.11937) | (LDPE method) | ICLR 2026 | 2024 | Countdown positional encoding achieves <3 token error for exact length control |
28
+ | 7 | [Dynamic Feedback for Length Regulation](https://arxiv.org/html/2601.01768) | — | arXiv | 2025 | Training-free dynamic feedback loop for length adherence |
29
+ | 8 | [Effective Context Engineering for AI Agents](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) | Anthropic Engineering | Blog | 2025 | Altitude calibration, finite attention budget, just-in-time context |
30
+ | 9 | [Control Illusion: Failure of Instruction Hierarchies](https://arxiv.org/pdf/2502.15851) | — | arXiv | 2025 | Instruction hierarchies break under conflicting constraints |
31
+ | 10 | [AgentIF: Instruction Following in Agentic Scenarios](https://arxiv.org/html/2505.16944v1) | — | arXiv | 2025 | Best models follow <30% of agentic instructions perfectly |
32
+ | 11 | [Persona Prompting as a Lens on LLM Social Reasoning](https://arxiv.org/abs/2601.20757) | — | arXiv | 2026 | Persona prompting improves classification but degrades rationale quality |
33
+
34
+ ### 2.2 Key Practitioner Sources
35
+
36
+ | Source | URL | Key Contribution |
37
+ |--------|-----|-----------------|
38
+ | Anthropic Context Engineering | [anthropic.com/engineering](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) | "Right altitude" system prompt design |
39
+ | Google Prompt Engineering Guide | [prompthub.us](https://www.prompthub.us/blog/googles-prompt-engineering-best-practices) | Positive framing over negation |
40
+ | DAIR.AI Prompt Engineering Guide | [promptingguide.ai](https://www.promptingguide.ai/papers) | Curated paper index |
41
+
42
+ ---
43
+
44
+ ## 3. Research Findings
45
+
46
+ ### 3.1 Prompt-Level Verbosity Control (Most Practical)
47
+
48
+ **Finding**: Natural language instructions in system prompts are the most effective and reliable mechanism for controlling response style. They categorically override activation-space interventions [Paper 4].
49
+
50
+ **Best practices distilled from literature**:
51
+
52
+ 1. **Positive framing over negation** — "Be concise" outperforms "Don't be verbose." KAIST research shows larger models perform worse on negated prompts [Paper 1, §2.2].
53
+
54
+ 2. **Explicit format/length constraints** — Specifying desired format, structure, length, and style yields the highest compliance. Example: "Respond in 2-3 sentences" vs. "Keep it short" [Paper 1].
55
+
56
+ 3. **Altitude calibration** (Anthropic) — System prompts should provide specific behavioral direction while remaining flexible. Neither too abstract ("be helpful") nor too brittle ("always respond in exactly 47 words") [Paper 8].
57
+
58
+ 4. **Adjective-based semantic anchoring** (SAC) — Grade behavioral intensity 1–5 using adjective clusters. For conciseness: Level 1 = "occasionally brief, somewhat terse"; Level 5 = "extremely concise, surgically precise, minimalist" [Paper 5].
59
+
60
+ 5. **Role/persona as behavioral anchor** — "You are a senior engineer reviewing code" naturally produces different verbosity than "You are a patient teacher explaining to a beginner" [Papers 1, 11].
61
+
62
+ ### 3.2 Length-Sensitive Prompt Design
63
+
64
+ **Finding**: LLM reasoning performance degrades at ~3,000 input tokens, with the practical sweet spot for most task prompts being 150–300 words [Paper 2].
65
+
66
+ **Finding**: U-shaped attention distribution means critical instructions should appear at the **beginning** or **end** of the system prompt, never buried in the middle [Paper 3].
67
+
68
+ **Implication for our system prompt**:
69
+ - Our SYSTEM_PROMPT is ~2,500 words (~3,500 tokens) — already near the degradation threshold
70
+ - Critical behavioral instructions (verbosity control, workflow) should be at the **top** and **bottom**
71
+ - Tool definitions (low-variance reference material) belong in the **middle** where attention is lowest
72
+
73
+ ### 3.3 Continuous Personality Dimensions (SAC Framework)
74
+
75
+ **Finding**: The SAC framework defines 5 behavioral intensity dimensions that can control any personality trait along a 1–5 scale [Paper 5]:
76
+
77
+ | Dimension | Controls | Verbosity Application |
78
+ |-----------|----------|----------------------|
79
+ | **Frequency** | How often the behavior occurs | How often the agent explains vs. acts silently |
80
+ | **Depth** | Emotional-cognitive engagement | How thoroughly the agent reasons in output |
81
+ | **Threshold** | Activation sensitivity | When the agent decides to narrate vs. stay silent |
82
+ | **Effort** | Energy invested in expression | How elaborate/polished the response is |
83
+ | **Willingness** | Voluntary commitment | How readily the agent offers unsolicited detail |
84
+
85
+ **Practical mapping for our agent**:
86
+
87
+ ```
88
+ Concise Mode (Level 1-2):
89
+ Frequency: "Rarely explain reasoning; act silently when possible"
90
+ Depth: "Surface-level status updates only"
91
+ Threshold: "Only speak when results are surprising or failed"
92
+ Effort: "Minimal formatting, no markdown headers for short answers"
93
+ Willingness: "Never volunteer extra context unless asked"
94
+
95
+ Verbose Mode (Level 4-5):
96
+ Frequency: "Explain each step and reasoning"
97
+ Depth: "Include technical details, alternatives considered"
98
+ Threshold: "Narrate even routine operations"
99
+ Effort: "Well-structured responses with headers and examples"
100
+ Willingness: "Proactively offer related context and suggestions"
101
+ ```
102
+
103
+ ### 3.4 Activation-Space Steering (Not Practical for Us)
104
+
105
+ **Finding**: Linear personality directions in activation space can probe personality traits but fail to steer behavior when explicit prompt instructions are present — "steering effects disappear entirely" when the prompt contains personality-relevant context [Paper 4].
106
+
107
+ **Implication**: This approach requires model weight access and fine-tuning. Not applicable to our inference-only framework (Ollama/vLLM). Prompt-level control is both more practical and more effective.
108
+
109
+ ### 3.5 Precise Length Control (LDPE)
110
+
111
+ **Finding**: Length-Difference Positional Encoding achieves <3 token error for exact length targeting [Paper 6]. However, this requires fine-tuning.
112
+
113
+ **Training-free alternative**: Dynamic feedback during generation can regulate length without fine-tuning [Paper 7]. This could be approximated at inference time through `max_tokens` parameter and system prompt length cues.
114
+
115
+ ### 3.6 Context Engineering for Multi-Turn Agents
116
+
117
+ **Key principles from Anthropic** [Paper 8]:
118
+
119
+ 1. **Finite attention budget** — Every added token depletes the model's attention capacity. Ruthlessly prioritize high-signal information.
120
+
121
+ 2. **Just-in-time context** — Don't pre-load; maintain lightweight identifiers and fetch dynamically.
122
+
123
+ 3. **Progressive disclosure** — Let the agent discover context through exploration rather than front-loading.
124
+
125
+ 4. **Compaction with recall** — Summarize while preserving critical details. Our Memex archive pattern already implements this.
126
+
127
+ 5. **Sub-agent architecture** — Deep exploration in child contexts, condensed summaries returned to parent. Already implemented in our sub_agent tool.
128
+
129
+ ---
130
+
131
+ ## 4. Integration Strategy for Omnius
132
+
133
+ ### 4.1 Transient Personality Injection Points
134
+
135
+ The agent framework has several natural injection points for transient style control:
136
+
137
+ | Injection Point | File | Mechanism |
138
+ |----------------|------|-----------|
139
+ | System prompt preamble | `agenticRunner.ts:223` | Static SYSTEM_PROMPT constant |
140
+ | Dynamic context injection | `agenticRunner.ts` `dynamicContext` option | Per-task context appended to system prompt |
141
+ | Health check prompts | `agenticRunner.ts:779` | Self-eval messages injected mid-task |
142
+ | Compaction summaries | `agenticRunner.ts:1728` | Style-control instructions preserved across compaction |
143
+ | Tool result formatting | `agenticRunner.ts:1819` | Observation masking controls output volume |
144
+
145
+ ### 4.2 Proposed: `PersonalityProfile` Interface
146
+
147
+ ```typescript
148
+ /**
149
+ * Transient personality profile for controlling agent response style.
150
+ * Based on SAC framework (arXiv:2506.20993) intensity dimensions.
151
+ *
152
+ * Each dimension is 1-5:
153
+ * 1 = minimal (concise, silent, terse)
154
+ * 3 = balanced (default)
155
+ * 5 = maximal (verbose, explanatory, thorough)
156
+ */
157
+ export interface PersonalityProfile {
158
+ /** How often the agent narrates its actions (1=silent, 5=running commentary) */
159
+ frequency: 1 | 2 | 3 | 4 | 5;
160
+ /** Depth of reasoning exposed in output (1=results only, 5=full chain of thought) */
161
+ depth: 1 | 2 | 3 | 4 | 5;
162
+ /** When the agent decides to speak vs. act silently (1=only on failure, 5=narrate everything) */
163
+ threshold: 1 | 2 | 3 | 4 | 5;
164
+ /** Effort invested in response formatting (1=raw, 5=polished markdown) */
165
+ effort: 1 | 2 | 3 | 4 | 5;
166
+ /** Willingness to offer unsolicited context (1=never, 5=proactive suggestions) */
167
+ willingness: 1 | 2 | 3 | 4 | 5;
168
+ }
169
+
170
+ /** Preset personality profiles */
171
+ export const PERSONALITY_PRESETS = {
172
+ /** Silent operator — acts, doesn't explain */
173
+ concise: { frequency: 1, depth: 1, threshold: 1, effort: 2, willingness: 1 },
174
+ /** Balanced default */
175
+ balanced: { frequency: 3, depth: 3, threshold: 3, effort: 3, willingness: 3 },
176
+ /** Thorough explainer — narrates reasoning */
177
+ verbose: { frequency: 5, depth: 4, threshold: 4, effort: 4, willingness: 4 },
178
+ /** Teacher mode — maximum explanation */
179
+ pedagogical: { frequency: 5, depth: 5, threshold: 5, effort: 5, willingness: 5 },
180
+ } as const;
181
+ ```
182
+
183
+ ### 4.3 Proposed: Personality-to-Prompt Compiler
184
+
185
+ The personality profile would compile to a system prompt suffix using adjective anchoring:
186
+
187
+ ```typescript
188
+ function compilePersonalityPrompt(profile: PersonalityProfile): string {
189
+ const avgIntensity = (profile.frequency + profile.depth +
190
+ profile.threshold + profile.effort + profile.willingness) / 5;
191
+
192
+ if (avgIntensity <= 1.5) {
193
+ return `\n## Response Style\nBe extremely concise. Act silently — only speak when results are ` +
194
+ `surprising or errors occur. No preamble, no summaries. Raw results and tool calls only.`;
195
+ }
196
+ if (avgIntensity <= 2.5) {
197
+ return `\n## Response Style\nBe concise and direct. Brief status updates between tool calls. ` +
198
+ `Skip reasoning explanation unless the approach is non-obvious. No markdown headers for short answers.`;
199
+ }
200
+ if (avgIntensity <= 3.5) {
201
+ return ``; // Default — no override needed
202
+ }
203
+ if (avgIntensity <= 4.5) {
204
+ return `\n## Response Style\nExplain your reasoning as you work. Describe what you're looking for ` +
205
+ `and why. Summarize findings. Use structured formatting for complex output.`;
206
+ }
207
+ return `\n## Response Style\nProvide thorough explanations of your reasoning at each step. ` +
208
+ `Describe alternatives you considered. Offer suggestions beyond the immediate task. ` +
209
+ `Use well-structured markdown with headers and examples.`;
210
+ }
211
+ ```
212
+
213
+ ### 4.4 System Prompt Restructuring (Position Optimization)
214
+
215
+ Based on "Lost in the Middle" findings [Paper 3], restructure the system prompt:
216
+
217
+ ```
218
+ [TOP — HIGH ATTENTION]
219
+ Identity + Core behavioral rules
220
+ Response style instructions (personality profile)
221
+ Critical workflow rules
222
+
223
+ [MIDDLE — LOW ATTENTION]
224
+ Tool definitions (reference material — rarely needs full attention)
225
+ Desktop automation details
226
+ Skills system
227
+
228
+ [BOTTOM — HIGH ATTENTION]
229
+ Critical Rules (validation, iteration, task_complete)
230
+ Project awareness
231
+ Dynamic context (memory, git state, project files)
232
+ ```
233
+
234
+ ### 4.5 Transient vs. Persistent Application
235
+
236
+ | Scope | Mechanism | When |
237
+ |-------|-----------|------|
238
+ | **Per-session** | User sets personality with `/style concise` command | User preference |
239
+ | **Per-task** | Task type detection adjusts profile (code review → verbose, quick fix → concise) | Automatic |
240
+ | **Per-turn** | Health check prompts can inject "be more/less verbose" | Self-correction |
241
+ | **Persistent** | Saved to `~/.omnius/config.json` as default preference | User configuration |
242
+
243
+ ### 4.6 Token Budget Implications
244
+
245
+ Per Levy et al. [Paper 2], our system prompt (~3,500 tokens) is near the reasoning degradation threshold. The personality prompt suffix should be:
246
+
247
+ - **Concise mode**: 0 extra tokens (suppress the suffix entirely)
248
+ - **Default mode**: 0 extra tokens (no suffix needed)
249
+ - **Verbose mode**: ~50 tokens max
250
+ - **Pedagogical mode**: ~80 tokens max
251
+
252
+ This keeps total system prompt overhead well within budget.
253
+
254
+ ---
255
+
256
+ ## 5. Implementation Roadmap
257
+
258
+ ### Phase 1: System Prompt Position Optimization
259
+ - Restructure SYSTEM_PROMPT per §4.4 (critical rules at top/bottom, tools in middle)
260
+ - Zero new features, pure prompt restructuring
261
+ - Verification: eval suite pass rate maintained or improved
262
+
263
+ ### Phase 2: PersonalityProfile Interface + Compiler
264
+ - Add PersonalityProfile type and preset profiles per §4.2
265
+ - Implement compilePersonalityPrompt() per §4.3
266
+ - Wire into AgenticRunner options
267
+ - Add `/style` slash command
268
+
269
+ ### Phase 3: Task-Adaptive Personality
270
+ - Detect task type (code review, bug fix, exploration, question answering)
271
+ - Auto-select personality profile based on task characteristics
272
+ - Allow user override
273
+
274
+ ### Phase 4: Self-Correcting Verbosity
275
+ - Health check prompts assess response style adherence
276
+ - Mid-task personality injection if agent drifts from target style
277
+ - Feedback loop with user satisfaction signal
278
+
279
+ ---
280
+
281
+ ## 6. References
282
+
283
+ 1. Schulhoff, S., et al. (2024). "The Prompt Report: A Systematic Survey of Prompting Techniques." arXiv:2406.06608. https://arxiv.org/abs/2406.06608
284
+ 2. Levy, M., Jacoby, A., & Goldberg, Y. (2024). "Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of LLMs." ACL 2024. https://arxiv.org/abs/2402.14848
285
+ 3. Liu, N.F., et al. (2024). "Lost in the Middle: How Language Models Use Long Contexts." TACL 12:157–173. https://arxiv.org/abs/2307.03172
286
+ 4. (2025). "Linear Personality Probing and Steering in LLMs: A Big Five Study." arXiv:2512.17639. https://arxiv.org/html/2512.17639v1
287
+ 5. (2026). "SAC: A Framework for Measuring and Inducing Personality Traits in LLMs with Dynamic Intensity Control." arXiv:2506.20993. https://arxiv.org/abs/2506.20993
288
+ 6. (2024). "Precise Length Control in Large Language Models." arXiv:2412.11937. ICLR 2026. https://arxiv.org/abs/2412.11937
289
+ 7. (2025). "Can LLMs Track Their Output Length? A Dynamic Feedback Mechanism for Precise Length Regulation." arXiv:2601.01768. https://arxiv.org/html/2601.01768
290
+ 8. Anthropic Engineering. (2025). "Effective Context Engineering for AI Agents." https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
291
+ 9. (2025). "Control Illusion: The Failure of Instruction Hierarchies." arXiv:2502.15851. https://arxiv.org/pdf/2502.15851
292
+ 10. (2025). "AgentIF: Benchmarking Instruction Following of LLMs in Agentic Scenarios." arXiv:2505.16944. https://arxiv.org/html/2505.16944v1
293
+ 11. (2026). "Persona Prompting as a Lens on LLM Social Reasoning." arXiv:2601.20757. https://arxiv.org/abs/2601.20757
@@ -2,6 +2,11 @@
2
2
 
3
3
  This directory is the human-readable REST API reference for Omnius. It is meant to be explored incrementally by agents and humans.
4
4
 
5
+ For intent-first lookup, start with
6
+ [Discovery](./endpoints/discovery.md), `GET /v1/discovery`, or the bundled
7
+ [`DISCOVERY.json`](../DISCOVERY.json). Use this index after the catalog points
8
+ to an endpoint family.
9
+
5
10
  Canonical sources:
6
11
 
7
12
  - Runtime OpenAPI generator: `packages/cli/src/api/openapi.ts`
@@ -21,6 +26,7 @@ Use this index first, then open the smallest relevant family file:
21
26
  | Auth, scopes, rate limits, API keys | `docs/rest/auth-and-scopes.md` |
22
27
  | Errors, pagination, ETags, request IDs | `docs/rest/errors-pagination-etags.md` |
23
28
  | OpenAPI and docs renderers | `docs/rest/openapi-source.md` |
29
+ | Capability and documentation discovery | `docs/rest/endpoints/discovery.md` |
24
30
  | Chat, realtime, OpenAI-compatible inference | `docs/rest/endpoints/chat.md` |
25
31
  | Agentic jobs and run lifecycle | `docs/rest/endpoints/run.md` |
26
32
  | Config, endpoints, keys, profiles, projects | `docs/rest/endpoints/config.md` |
@@ -88,6 +94,7 @@ Scopes:
88
94
  | Family | Representative endpoints |
89
95
  | --- | --- |
90
96
  | Health | `/health`, `/health/ready`, `/health/startup`, `/version`, `/metrics` |
97
+ | Discovery | `/v1/discovery`, `/v1/discovery/{id}` |
91
98
  | Inference and chat | `/v1/models`, `/v1/chat/completions`, `/v1/embeddings`, `/v1/chat`, `/v1/generate`, `/api/generate`, `/v1/chat/sessions`, `/v1/chat/check-in` |
92
99
  | AIWG | `/v1/aiwg`, `/v1/aiwg/frameworks`, `/v1/aiwg/skills`, `/v1/aiwg/use`, `/v1/aiwg/expand` |
93
100
  | Runs | `/v1/run`, `/v1/runs`, `/v1/runs/{id}`, `/v1/todos`, `/v1/todos/{session_id}` |
@@ -41,6 +41,13 @@ curl -s http://127.0.0.1:11435/health/ready
41
41
  curl -s http://127.0.0.1:11435/version
42
42
  ```
43
43
 
44
+ ## Discovery
45
+
46
+ ```bash
47
+ curl -s 'http://127.0.0.1:11435/v1/discovery?q=bring%20your%20own%20inference'
48
+ curl -s http://127.0.0.1:11435/v1/discovery/tool.web-search
49
+ ```
50
+
44
51
  ## Models
45
52
 
46
53
  ```bash
@@ -104,12 +111,23 @@ curl -s http://127.0.0.1:11435/v1/skills/omnius-rest-docs
104
111
 
105
112
  ## Tool Call
106
113
 
114
+ First inspect `direct_callable`:
115
+
116
+ ```bash
117
+ curl -s http://127.0.0.1:11435/v1/tools/memory_search
118
+ ```
119
+
107
120
  ```bash
108
121
  curl -s -X POST http://127.0.0.1:11435/v1/tools/memory_search/call \
109
122
  -H 'content-type: application/json' \
110
123
  -d '{"args":{"query":"rest api docs"}}'
111
124
  ```
112
125
 
126
+ `web_search` is agent-bound. Offer it to `/v1/run`, `/v1/chat`, or
127
+ `/v1/chat/completions` with `agent_loop: true` and
128
+ `include_daemon_tools: ["read"]`; do not infer direct-call support from the
129
+ presence of its schema.
130
+
113
131
  Deterministic bookkeeping without an agent run:
114
132
 
115
133
  ```bash
@@ -12,6 +12,7 @@
12
12
  "full_reference": "docs/reference/rest-api.md",
13
13
  "skill": ".aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md",
14
14
  "families": [
15
+ { "name": "discovery", "path": "docs/rest/endpoints/discovery.md", "triggers": ["discovery", "capabilities", "catalog", "/v1/discovery", "discover", "show"] },
15
16
  { "name": "chat", "path": "docs/rest/endpoints/chat.md", "triggers": ["chat", "generate", "realtime", "OpenAI-compatible", "/v1/chat", "/v1/generate"] },
16
17
  { "name": "run", "path": "docs/rest/endpoints/run.md", "triggers": ["run", "jobs", "todos", "scheduled"] },
17
18
  { "name": "config", "path": "docs/rest/endpoints/config.md", "triggers": ["config", "keys", "profiles", "projects", "endpoint"] },
@@ -24,7 +24,13 @@ key:scope:owner:rpm:tpd:max_jobs
24
24
 
25
25
  Compatibility: `OMNIUS_API_KEY` and `OMNIUS_API_KEYS` are still accepted as legacy REST auth variables, but new deployments should use the `OMNIUS_REST_*` namespace.
26
26
 
27
- Provider keys are separate from REST keys. Prefer `OMNIUS_PROVIDER_API_KEY`, `OMNIUS_MODEL_API_KEY`, or provider-native variables such as `VLLM_API_KEY` for upstream model auth. `OMNIUS_API_KEY` remains a legacy fallback for provider auth and legacy REST auth, so avoid it in new remote deployments. Scoped REST keys (`OMNIUS_REST_*` and `OMNIUS_API_KEYS`) are stripped before spawning agent subprocesses and are not forwarded as model-provider credentials.
27
+ Provider keys are separate from REST keys. The exact upstream precedence is
28
+ `OMNIUS_PROVIDER_API_KEY` → `OMNIUS_MODEL_API_KEY` →
29
+ `OMNIUS_UPSTREAM_API_KEY` → `OMNIUS_API_KEY` → `VLLM_API_KEY` → persisted
30
+ endpoint configuration. `OMNIUS_API_KEY` remains a legacy fallback for both
31
+ provider and legacy REST auth, so avoid it in new remote deployments. Scoped
32
+ REST keys (`OMNIUS_REST_*` and `OMNIUS_API_KEYS`) are stripped before spawning
33
+ agent subprocesses and are not forwarded as model-provider credentials.
28
34
 
29
35
  Fields:
30
36
 
@@ -0,0 +1,44 @@
1
+ # Discovery
2
+
3
+ The discovery API exposes the same catalog shipped in `docs/DISCOVERY.json`
4
+ and used by `omnius discover` / `omnius show`.
5
+
6
+ | Method | Path | Purpose |
7
+ | --- | --- | --- |
8
+ | `GET` | `/v1/discovery` | Search or list catalog entries |
9
+ | `GET` | `/v1/discovery/{id}` | Expand one stable entry |
10
+
11
+ ## Search
12
+
13
+ ```bash
14
+ curl -s "http://127.0.0.1:11435/v1/discovery?q=web%20search&kind=tool&limit=5"
15
+ ```
16
+
17
+ Query fields:
18
+
19
+ - `q`: free-text intent; omit to list entries.
20
+ - `kind`: one supported catalog kind.
21
+ - `limit`: page size.
22
+ - `offset`: zero-based page offset.
23
+
24
+ The response includes pagination metadata and an ETag. Use
25
+ `If-None-Match` when polling a long-running daemon.
26
+
27
+ ## Expand
28
+
29
+ ```bash
30
+ curl -s http://127.0.0.1:11435/v1/discovery/tool.web-search
31
+ ```
32
+
33
+ An entry contains its stable ID, kind, title, summary, aliases/keywords,
34
+ invocation interfaces, typed references, and related entries. Discovery is
35
+ read-scoped and does not execute the selected capability.
36
+
37
+ Important entrypoints:
38
+
39
+ - `capability.bring-your-own-inference`
40
+ - `provider.anthropic`
41
+ - `provider.gemini`
42
+ - `tool.web-search`
43
+ - `api.tools`
44
+ - `operation.version-compatibility`
@@ -61,3 +61,8 @@ Usage endpoints support dashboards for requests, token counts, rate limits, and
61
61
  ## Ollama Pool
62
62
 
63
63
  The pool endpoints are admin-oriented process hygiene tools. They report managed `ollama serve` processes and cleanup decisions. Cleanup accounts for stale pool state and orphan runner risks.
64
+
65
+ `GET /version` is also the compatibility bootstrap for an agent or service
66
+ integration. Execution clients should send `X-Omnius-Min-Version` on work
67
+ requests so an older daemon returns `412` before starting work. See
68
+ [Service Version Compatibility](../../operations/version-compatibility.md).
@@ -61,6 +61,15 @@ Profiles can also be selected with `X-Tool-Profile`. Named profiles resolve in t
61
61
 
62
62
  Tool calls are gated by auth scope, tool policy, off-device policy, and profile restrictions.
63
63
 
64
+ Catalog presence does not imply direct invocation. Inspect
65
+ `direct_callable` in the live metadata. In particular, `web_search` is
66
+ agent-bound: discover its schema with `GET /v1/tools/web_search`, then offer it
67
+ through `/v1/run`, `/v1/chat`, or `/v1/chat/completions` using
68
+ `agent_loop: true` and `include_daemon_tools: ["read"]`. Do not assume
69
+ `POST /v1/tools/web_search/call` is available.
70
+
71
+ See [Tools And Web Search](../../guides/tools-and-web-search.md).
72
+
64
73
  Bookkeeping tools are direct-callable and do not require a model run:
65
74
 
66
75
  ```bash
@@ -0,0 +1,42 @@
1
+ # Adversary System Review
2
+
3
+ **Date:** 2026-06-28
4
+ **Scope:** Read-only pass
5
+ **Files reviewed:**
6
+ - `packages/orchestrator/src/adversaryStream.ts` (407 lines)
7
+ - `packages/orchestrator/tests/adversaryStream.test.ts` (242 lines)
8
+
9
+ ## Architecture
10
+
11
+ The adversary is a **generative, inference-driven auditor** running as a separate memory stream (`AdversaryStream`) adjacent to the main agent loop. It replaces the prior pure-regex adversary that only fired `false_success` when there were ZERO recent successes (missing every mixed success/failure run).
12
+
13
+ ## Key Features
14
+
15
+ 1. **Own persistent memory** — claims ledger (assertions, evidence state, verdicts) separate from main agent context
16
+ 2. **Single inference per observation** — fires one backend call per high-signal claim, parses structured critique
17
+ 3. **Classification taxonomy** — `false_failure`, `redundant_action`, `loop_detected`, plus `ok` (no intervention)
18
+ 4. **Loop detection** — `loopSignal` tracks repeated tool calls with same target; fires when count ≥ threshold
19
+ 5. **Success-language regex** — `SUCCESS_LANGUAGE` matches optimistic claims like "done", "verified", "live", "hydrated"
20
+ 6. **Non-blocking** — adversary is advisory; never blocks tool calls. `adversaryRedundantSignal` is Priority 1 in critic but returns guidance, not a block.
21
+
22
+ ## Integration Points
23
+
24
+ - **AgenticRunner**: manages `adversaryMode` ("backseat" | "skillcoach" | "both"), creates `AdversaryStream`, emits `debug_adversary` and `adversary_reaction` events
25
+ - **Critic**: `adversaryRedundantSignal` triggers WO-FIX-C path with message "The adversary recognized this exact tool call as already observed earlier."
26
+ - **TUI**: renders adversary buffer (last 50 entries), shows `adversary: N tracked` in status line
27
+
28
+ ## Test Coverage
29
+
30
+ - `adversarySystemPrompt` encodes evidence-demanding posture (checks "started ≠ running", "exit code 0")
31
+ - `parseAdversaryCritique` structured output parsing
32
+ - `AdversaryStream` single-flight behavior, ledger persistence
33
+ - Loop detection with `loopSignal`
34
+ - `false_failure` and `redundant_action` classification
35
+
36
+ ## Assessment
37
+
38
+ The system is well-structured. The generative approach (inference-driven vs regex) addresses the original blind spot. The non-blocking design is correct — adversary provides critique without preventing progress. The integration with critic's Priority 1 path gives it real influence on loop behavior.
39
+
40
+ ## No Issues Found
41
+
42
+ This was a read-only review. No bugs, no improvements identified. The system is production-ready.