omnius 1.0.591 → 1.0.592
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.aiwg/addons/omnius-docs/README.md +15 -1
- package/.aiwg/addons/omnius-docs/manifest.json +28 -68
- package/.aiwg/addons/omnius-docs/skills/agent-failure-recovery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/browser-interaction-validation/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/evidence-directed-delivery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/hardware-evidence-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/omnius-docs/SKILL.md +17 -7
- package/.aiwg/addons/omnius-docs/skills/omnius-inference-docs/SKILL.md +27 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-integration-docs/SKILL.md +21 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-ops-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-realtime-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-sponsor-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-telegram-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-tools-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-version-compatibility-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/runtime-provenance-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/secrets-and-config-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/test-surface-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/workspace-reality-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-rest-docs/README.md +3 -0
- package/.aiwg/addons/omnius-rest-docs/manifest.json +27 -20
- package/.aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md +9 -5
- package/README.md +36 -0
- package/dist/discovery.d.ts +50 -0
- package/dist/index.js +5975 -4021
- package/dist/library.d.ts +7 -0
- package/dist/library.js +950 -0
- package/dist/postinstall-daemon.cjs +18 -0
- package/dist/providerRegistry.d.ts +80 -0
- package/dist/service-version.d.ts +35 -0
- package/docs/.vitepress/config.mts +8 -0
- package/docs/DISCOVERY.json +20224 -0
- package/docs/DISCOVERY.md +648 -0
- package/docs/HANDOFF-crl-encoder-decoder-fix.md +129 -0
- package/docs/agent-memory/INDEX.md +9 -4
- package/docs/agent-memory/index.md +7 -0
- package/docs/concept-relational-language.md +869 -0
- package/docs/context-management-medium-models-proposal.md +449 -0
- package/docs/dedup-false-positive-meta-analysis.md +96 -0
- package/docs/discovery/catalog-overrides.json +724 -0
- package/docs/duplicate-calls-root-cause-analysis.md +91 -0
- package/docs/duplicate-calls-root-cause-deep.md +155 -0
- package/docs/ephemeral-skill-pack-small-context.md +57 -0
- package/docs/explorations/context-window-todo-association.md +156 -0
- package/docs/explorations/todo-association-verify.json +30 -0
- package/docs/explorations/verification-ledger.json +45 -0
- package/docs/explorations/verify-todo-association.sh +30 -0
- package/docs/flowstate.md +806 -0
- package/docs/getting-started/install.md +24 -0
- package/docs/getting-started/model-providers.md +13 -0
- package/docs/guides/agent-integration.md +87 -0
- package/docs/guides/bring-your-own-inference.md +126 -0
- package/docs/guides/tools-and-web-search.md +95 -0
- package/docs/index.md +14 -0
- package/docs/longhaul-35b-workorders.md +496 -0
- package/docs/memory-integration-analysis.md +303 -0
- package/docs/model-capability-awareness-and-multimodal-memory-root-fix.md +799 -0
- package/docs/multimodal-identity-memory-implementation.md +76 -0
- package/docs/omnius-self-edit-eval-2026-06-10.md +169 -0
- package/docs/opencode-agentic-loop-comparison.md +290 -0
- package/docs/operations/security-and-remote-access.md +2 -2
- package/docs/operations/version-compatibility.md +63 -0
- package/docs/proposals/git-progress-tracking-strategy.md +289 -0
- package/docs/proposals/opencode-modules/backendAdapter.ts +443 -0
- package/docs/proposals/opencode-modules/childSession.ts +288 -0
- package/docs/proposals/opencode-modules/compactionAgent.ts +101 -0
- package/docs/proposals/opencode-modules/orchestrator.ts +387 -0
- package/docs/proposals/opencode-modules/runner.ts +258 -0
- package/docs/reference/auth-map.md +87 -196
- package/docs/reference/configuration.md +27 -0
- package/docs/reference/rest-api.md +7 -0
- package/docs/reference/slash-commands.md +125 -2
- package/docs/research/_archived/README.md +18 -0
- package/docs/research/_archived/context_window_attention_model.py +418 -0
- package/docs/research/_archived/context_window_attention_spec.md +55 -0
- package/docs/research/_archived/context_window_attention_weights.json +68 -0
- package/docs/research/k-splanifolds.pdf +0 -0
- package/docs/research/personality-verbosity-control.md +293 -0
- package/docs/rest/INDEX.md +7 -0
- package/docs/rest/QUICKREF.md +18 -0
- package/docs/rest/REST-DOCS-MANIFEST.json +1 -0
- package/docs/rest/auth-and-scopes.md +7 -1
- package/docs/rest/endpoints/discovery.md +44 -0
- package/docs/rest/endpoints/events.md +5 -0
- package/docs/rest/endpoints/tools.md +9 -0
- package/docs/reviews/adversary-system-review.md +42 -0
- package/docs/sana-and-video-generation-integration-plan.md +712 -0
- package/docs/session-diary-llm-training-analysis.md +218 -0
- package/docs/telegram-dmn-curiosity-outreach-scaffold.md +91 -0
- package/docs/telegram-mid-horizon-download-loop-handoff.md +468 -0
- package/docs/telegram-reflection-corpus-integration-plan.md +306 -0
- package/docs/telegram-unified-tooling-architecture.md +332 -0
- package/docs/threat-model.md +868 -0
- package/docs/trajectory-grounding.md +160 -0
- package/docs/voice-flow-architecture.md +489 -0
- package/docs/work-orders/WO-AM-GAPS.md +638 -0
- package/docs/work-orders/daemon-hud-ui-overhaul.md +82 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/INDEX.md +21 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/WORKORDER.md +225 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/WORKORDER.md +198 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/WORKORDER.md +172 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/WORKORDER.md +169 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/WORKORDER.md +189 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/WORKORDER.md +199 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/WORKORDER.md +174 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/WORKORDER.md +226 -0
- package/docs/work-orders/hermes-architecture-deltas/INDEX.md +38 -0
- package/docs/work-orders/omnius-context-engineering-behavior-fixes.md +281 -0
- package/docs/work-orders/telegram-dropbear-context-rca-workorder.md +202 -0
- package/docs/work-orders/world-class-memory-compiler/README.md +162 -0
- package/docs/work-orders/world-class-memory-compiler/TRACKER.md +179 -0
- package/docs/work-orders/world-class-memory-compiler/WO-01-exact-request-budget.md +79 -0
- package/docs/work-orders/world-class-memory-compiler/WO-02-typed-memory-fabric.md +65 -0
- package/docs/work-orders/world-class-memory-compiler/WO-03-dependency-working-set.md +55 -0
- package/docs/work-orders/world-class-memory-compiler/WO-04-inference-memory-compiler.md +67 -0
- package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md +72 -0
- package/docs/work-orders/world-class-memory-compiler/WO-06-temporal-hybrid-retrieval.md +49 -0
- package/docs/work-orders/world-class-memory-compiler/WO-07-evaluation-harness.md +45 -0
- package/docs/work-orders/world-class-memory-compiler/WO-08-rollout-legacy-removal.md +45 -0
- package/docs/x402-remote-inference-plan.md +323 -0
- package/npm-shrinkwrap.json +108 -117
- package/package.json +7 -6
- package/templates/AGENTS.md +6 -0
- package/templates/OMNIUS.md +20 -0
|
@@ -0,0 +1,293 @@
|
|
|
1
|
+
# LLM Personality Guidance & Verbosity Control: Research Synthesis
|
|
2
|
+
|
|
3
|
+
**Date**: 2026-03-13
|
|
4
|
+
**Purpose**: Literature review and integration strategy for transient personality/verbosity steering in the Omnius framework
|
|
5
|
+
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
## 1. Executive Summary
|
|
9
|
+
|
|
10
|
+
This document synthesizes research on controlling LLM output style — specifically verbosity vs. conciseness — through system prompt engineering, persona/personality steering, and length control mechanisms. The findings are organized for direct integration into the Omnius agentic framework, where the system prompt and context engineering pipeline can transiently influence agent behavior without fine-tuning.
|
|
11
|
+
|
|
12
|
+
**Key insight**: Prompt-level personality steering (natural language instructions) categorically dominates activation-level interventions. Explicit system prompt instructions override all other behavioral signals, making prompt engineering the most practical lever for our framework.
|
|
13
|
+
|
|
14
|
+
---
|
|
15
|
+
|
|
16
|
+
## 2. Literature & Provenance
|
|
17
|
+
|
|
18
|
+
### 2.1 Core Papers
|
|
19
|
+
|
|
20
|
+
| # | Paper | Authors | Venue | Year | Key Finding |
|
|
21
|
+
|---|-------|---------|-------|------|-------------|
|
|
22
|
+
| 1 | [The Prompt Report: A Systematic Survey of Prompting Techniques](https://arxiv.org/abs/2406.06608) | Schulhoff et al. (32 authors) | arXiv | 2024 | Taxonomy of 58 prompting techniques; role/persona prompting in §2.2 |
|
|
23
|
+
| 2 | [Same Task, More Tokens](https://arxiv.org/abs/2402.14848) | Levy, Jacoby & Goldberg | ACL 2024 | 2024 | LLM reasoning degrades ~3,000 tokens; sweet spot 150–300 words |
|
|
24
|
+
| 3 | [Lost in the Middle](https://arxiv.org/abs/2307.03172) | Liu, Lin, Hewitt et al. | TACL 2024 | 2024 | U-shaped attention bias; >30% accuracy drop for mid-context info |
|
|
25
|
+
| 4 | [Linear Personality Probing and Steering in LLMs](https://arxiv.org/html/2512.17639v1) | (Big Five study) | arXiv | 2025 | Linear directions can probe personality but explicit prompts override steering vectors entirely |
|
|
26
|
+
| 5 | [SAC: Style Adjective Continuation](https://arxiv.org/abs/2506.20993) | (16PF framework) | arXiv | 2026 | Continuous 1–5 trait intensity via adjective anchoring; 5 behavioral dimensions |
|
|
27
|
+
| 6 | [Precise Length Control in LLMs (LDPE)](https://arxiv.org/abs/2412.11937) | (LDPE method) | ICLR 2026 | 2024 | Countdown positional encoding achieves <3 token error for exact length control |
|
|
28
|
+
| 7 | [Dynamic Feedback for Length Regulation](https://arxiv.org/html/2601.01768) | — | arXiv | 2025 | Training-free dynamic feedback loop for length adherence |
|
|
29
|
+
| 8 | [Effective Context Engineering for AI Agents](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) | Anthropic Engineering | Blog | 2025 | Altitude calibration, finite attention budget, just-in-time context |
|
|
30
|
+
| 9 | [Control Illusion: Failure of Instruction Hierarchies](https://arxiv.org/pdf/2502.15851) | — | arXiv | 2025 | Instruction hierarchies break under conflicting constraints |
|
|
31
|
+
| 10 | [AgentIF: Instruction Following in Agentic Scenarios](https://arxiv.org/html/2505.16944v1) | — | arXiv | 2025 | Best models follow <30% of agentic instructions perfectly |
|
|
32
|
+
| 11 | [Persona Prompting as a Lens on LLM Social Reasoning](https://arxiv.org/abs/2601.20757) | — | arXiv | 2026 | Persona prompting improves classification but degrades rationale quality |
|
|
33
|
+
|
|
34
|
+
### 2.2 Key Practitioner Sources
|
|
35
|
+
|
|
36
|
+
| Source | URL | Key Contribution |
|
|
37
|
+
|--------|-----|-----------------|
|
|
38
|
+
| Anthropic Context Engineering | [anthropic.com/engineering](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) | "Right altitude" system prompt design |
|
|
39
|
+
| Google Prompt Engineering Guide | [prompthub.us](https://www.prompthub.us/blog/googles-prompt-engineering-best-practices) | Positive framing over negation |
|
|
40
|
+
| DAIR.AI Prompt Engineering Guide | [promptingguide.ai](https://www.promptingguide.ai/papers) | Curated paper index |
|
|
41
|
+
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
## 3. Research Findings
|
|
45
|
+
|
|
46
|
+
### 3.1 Prompt-Level Verbosity Control (Most Practical)
|
|
47
|
+
|
|
48
|
+
**Finding**: Natural language instructions in system prompts are the most effective and reliable mechanism for controlling response style. They categorically override activation-space interventions [Paper 4].
|
|
49
|
+
|
|
50
|
+
**Best practices distilled from literature**:
|
|
51
|
+
|
|
52
|
+
1. **Positive framing over negation** — "Be concise" outperforms "Don't be verbose." KAIST research shows larger models perform worse on negated prompts [Paper 1, §2.2].
|
|
53
|
+
|
|
54
|
+
2. **Explicit format/length constraints** — Specifying desired format, structure, length, and style yields the highest compliance. Example: "Respond in 2-3 sentences" vs. "Keep it short" [Paper 1].
|
|
55
|
+
|
|
56
|
+
3. **Altitude calibration** (Anthropic) — System prompts should provide specific behavioral direction while remaining flexible. Neither too abstract ("be helpful") nor too brittle ("always respond in exactly 47 words") [Paper 8].
|
|
57
|
+
|
|
58
|
+
4. **Adjective-based semantic anchoring** (SAC) — Grade behavioral intensity 1–5 using adjective clusters. For conciseness: Level 1 = "occasionally brief, somewhat terse"; Level 5 = "extremely concise, surgically precise, minimalist" [Paper 5].
|
|
59
|
+
|
|
60
|
+
5. **Role/persona as behavioral anchor** — "You are a senior engineer reviewing code" naturally produces different verbosity than "You are a patient teacher explaining to a beginner" [Papers 1, 11].
|
|
61
|
+
|
|
62
|
+
### 3.2 Length-Sensitive Prompt Design
|
|
63
|
+
|
|
64
|
+
**Finding**: LLM reasoning performance degrades at ~3,000 input tokens, with the practical sweet spot for most task prompts being 150–300 words [Paper 2].
|
|
65
|
+
|
|
66
|
+
**Finding**: U-shaped attention distribution means critical instructions should appear at the **beginning** or **end** of the system prompt, never buried in the middle [Paper 3].
|
|
67
|
+
|
|
68
|
+
**Implication for our system prompt**:
|
|
69
|
+
- Our SYSTEM_PROMPT is ~2,500 words (~3,500 tokens) — already near the degradation threshold
|
|
70
|
+
- Critical behavioral instructions (verbosity control, workflow) should be at the **top** and **bottom**
|
|
71
|
+
- Tool definitions (low-variance reference material) belong in the **middle** where attention is lowest
|
|
72
|
+
|
|
73
|
+
### 3.3 Continuous Personality Dimensions (SAC Framework)
|
|
74
|
+
|
|
75
|
+
**Finding**: The SAC framework defines 5 behavioral intensity dimensions that can control any personality trait along a 1–5 scale [Paper 5]:
|
|
76
|
+
|
|
77
|
+
| Dimension | Controls | Verbosity Application |
|
|
78
|
+
|-----------|----------|----------------------|
|
|
79
|
+
| **Frequency** | How often the behavior occurs | How often the agent explains vs. acts silently |
|
|
80
|
+
| **Depth** | Emotional-cognitive engagement | How thoroughly the agent reasons in output |
|
|
81
|
+
| **Threshold** | Activation sensitivity | When the agent decides to narrate vs. stay silent |
|
|
82
|
+
| **Effort** | Energy invested in expression | How elaborate/polished the response is |
|
|
83
|
+
| **Willingness** | Voluntary commitment | How readily the agent offers unsolicited detail |
|
|
84
|
+
|
|
85
|
+
**Practical mapping for our agent**:
|
|
86
|
+
|
|
87
|
+
```
|
|
88
|
+
Concise Mode (Level 1-2):
|
|
89
|
+
Frequency: "Rarely explain reasoning; act silently when possible"
|
|
90
|
+
Depth: "Surface-level status updates only"
|
|
91
|
+
Threshold: "Only speak when results are surprising or failed"
|
|
92
|
+
Effort: "Minimal formatting, no markdown headers for short answers"
|
|
93
|
+
Willingness: "Never volunteer extra context unless asked"
|
|
94
|
+
|
|
95
|
+
Verbose Mode (Level 4-5):
|
|
96
|
+
Frequency: "Explain each step and reasoning"
|
|
97
|
+
Depth: "Include technical details, alternatives considered"
|
|
98
|
+
Threshold: "Narrate even routine operations"
|
|
99
|
+
Effort: "Well-structured responses with headers and examples"
|
|
100
|
+
Willingness: "Proactively offer related context and suggestions"
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
### 3.4 Activation-Space Steering (Not Practical for Us)
|
|
104
|
+
|
|
105
|
+
**Finding**: Linear personality directions in activation space can probe personality traits but fail to steer behavior when explicit prompt instructions are present — "steering effects disappear entirely" when the prompt contains personality-relevant context [Paper 4].
|
|
106
|
+
|
|
107
|
+
**Implication**: This approach requires model weight access and fine-tuning. Not applicable to our inference-only framework (Ollama/vLLM). Prompt-level control is both more practical and more effective.
|
|
108
|
+
|
|
109
|
+
### 3.5 Precise Length Control (LDPE)
|
|
110
|
+
|
|
111
|
+
**Finding**: Length-Difference Positional Encoding achieves <3 token error for exact length targeting [Paper 6]. However, this requires fine-tuning.
|
|
112
|
+
|
|
113
|
+
**Training-free alternative**: Dynamic feedback during generation can regulate length without fine-tuning [Paper 7]. This could be approximated at inference time through `max_tokens` parameter and system prompt length cues.
|
|
114
|
+
|
|
115
|
+
### 3.6 Context Engineering for Multi-Turn Agents
|
|
116
|
+
|
|
117
|
+
**Key principles from Anthropic** [Paper 8]:
|
|
118
|
+
|
|
119
|
+
1. **Finite attention budget** — Every added token depletes the model's attention capacity. Ruthlessly prioritize high-signal information.
|
|
120
|
+
|
|
121
|
+
2. **Just-in-time context** — Don't pre-load; maintain lightweight identifiers and fetch dynamically.
|
|
122
|
+
|
|
123
|
+
3. **Progressive disclosure** — Let the agent discover context through exploration rather than front-loading.
|
|
124
|
+
|
|
125
|
+
4. **Compaction with recall** — Summarize while preserving critical details. Our Memex archive pattern already implements this.
|
|
126
|
+
|
|
127
|
+
5. **Sub-agent architecture** — Deep exploration in child contexts, condensed summaries returned to parent. Already implemented in our sub_agent tool.
|
|
128
|
+
|
|
129
|
+
---
|
|
130
|
+
|
|
131
|
+
## 4. Integration Strategy for Omnius
|
|
132
|
+
|
|
133
|
+
### 4.1 Transient Personality Injection Points
|
|
134
|
+
|
|
135
|
+
The agent framework has several natural injection points for transient style control:
|
|
136
|
+
|
|
137
|
+
| Injection Point | File | Mechanism |
|
|
138
|
+
|----------------|------|-----------|
|
|
139
|
+
| System prompt preamble | `agenticRunner.ts:223` | Static SYSTEM_PROMPT constant |
|
|
140
|
+
| Dynamic context injection | `agenticRunner.ts` `dynamicContext` option | Per-task context appended to system prompt |
|
|
141
|
+
| Health check prompts | `agenticRunner.ts:779` | Self-eval messages injected mid-task |
|
|
142
|
+
| Compaction summaries | `agenticRunner.ts:1728` | Style-control instructions preserved across compaction |
|
|
143
|
+
| Tool result formatting | `agenticRunner.ts:1819` | Observation masking controls output volume |
|
|
144
|
+
|
|
145
|
+
### 4.2 Proposed: `PersonalityProfile` Interface
|
|
146
|
+
|
|
147
|
+
```typescript
|
|
148
|
+
/**
|
|
149
|
+
* Transient personality profile for controlling agent response style.
|
|
150
|
+
* Based on SAC framework (arXiv:2506.20993) intensity dimensions.
|
|
151
|
+
*
|
|
152
|
+
* Each dimension is 1-5:
|
|
153
|
+
* 1 = minimal (concise, silent, terse)
|
|
154
|
+
* 3 = balanced (default)
|
|
155
|
+
* 5 = maximal (verbose, explanatory, thorough)
|
|
156
|
+
*/
|
|
157
|
+
export interface PersonalityProfile {
|
|
158
|
+
/** How often the agent narrates its actions (1=silent, 5=running commentary) */
|
|
159
|
+
frequency: 1 | 2 | 3 | 4 | 5;
|
|
160
|
+
/** Depth of reasoning exposed in output (1=results only, 5=full chain of thought) */
|
|
161
|
+
depth: 1 | 2 | 3 | 4 | 5;
|
|
162
|
+
/** When the agent decides to speak vs. act silently (1=only on failure, 5=narrate everything) */
|
|
163
|
+
threshold: 1 | 2 | 3 | 4 | 5;
|
|
164
|
+
/** Effort invested in response formatting (1=raw, 5=polished markdown) */
|
|
165
|
+
effort: 1 | 2 | 3 | 4 | 5;
|
|
166
|
+
/** Willingness to offer unsolicited context (1=never, 5=proactive suggestions) */
|
|
167
|
+
willingness: 1 | 2 | 3 | 4 | 5;
|
|
168
|
+
}
|
|
169
|
+
|
|
170
|
+
/** Preset personality profiles */
|
|
171
|
+
export const PERSONALITY_PRESETS = {
|
|
172
|
+
/** Silent operator — acts, doesn't explain */
|
|
173
|
+
concise: { frequency: 1, depth: 1, threshold: 1, effort: 2, willingness: 1 },
|
|
174
|
+
/** Balanced default */
|
|
175
|
+
balanced: { frequency: 3, depth: 3, threshold: 3, effort: 3, willingness: 3 },
|
|
176
|
+
/** Thorough explainer — narrates reasoning */
|
|
177
|
+
verbose: { frequency: 5, depth: 4, threshold: 4, effort: 4, willingness: 4 },
|
|
178
|
+
/** Teacher mode — maximum explanation */
|
|
179
|
+
pedagogical: { frequency: 5, depth: 5, threshold: 5, effort: 5, willingness: 5 },
|
|
180
|
+
} as const;
|
|
181
|
+
```
|
|
182
|
+
|
|
183
|
+
### 4.3 Proposed: Personality-to-Prompt Compiler
|
|
184
|
+
|
|
185
|
+
The personality profile would compile to a system prompt suffix using adjective anchoring:
|
|
186
|
+
|
|
187
|
+
```typescript
|
|
188
|
+
function compilePersonalityPrompt(profile: PersonalityProfile): string {
|
|
189
|
+
const avgIntensity = (profile.frequency + profile.depth +
|
|
190
|
+
profile.threshold + profile.effort + profile.willingness) / 5;
|
|
191
|
+
|
|
192
|
+
if (avgIntensity <= 1.5) {
|
|
193
|
+
return `\n## Response Style\nBe extremely concise. Act silently — only speak when results are ` +
|
|
194
|
+
`surprising or errors occur. No preamble, no summaries. Raw results and tool calls only.`;
|
|
195
|
+
}
|
|
196
|
+
if (avgIntensity <= 2.5) {
|
|
197
|
+
return `\n## Response Style\nBe concise and direct. Brief status updates between tool calls. ` +
|
|
198
|
+
`Skip reasoning explanation unless the approach is non-obvious. No markdown headers for short answers.`;
|
|
199
|
+
}
|
|
200
|
+
if (avgIntensity <= 3.5) {
|
|
201
|
+
return ``; // Default — no override needed
|
|
202
|
+
}
|
|
203
|
+
if (avgIntensity <= 4.5) {
|
|
204
|
+
return `\n## Response Style\nExplain your reasoning as you work. Describe what you're looking for ` +
|
|
205
|
+
`and why. Summarize findings. Use structured formatting for complex output.`;
|
|
206
|
+
}
|
|
207
|
+
return `\n## Response Style\nProvide thorough explanations of your reasoning at each step. ` +
|
|
208
|
+
`Describe alternatives you considered. Offer suggestions beyond the immediate task. ` +
|
|
209
|
+
`Use well-structured markdown with headers and examples.`;
|
|
210
|
+
}
|
|
211
|
+
```
|
|
212
|
+
|
|
213
|
+
### 4.4 System Prompt Restructuring (Position Optimization)
|
|
214
|
+
|
|
215
|
+
Based on "Lost in the Middle" findings [Paper 3], restructure the system prompt:
|
|
216
|
+
|
|
217
|
+
```
|
|
218
|
+
[TOP — HIGH ATTENTION]
|
|
219
|
+
Identity + Core behavioral rules
|
|
220
|
+
Response style instructions (personality profile)
|
|
221
|
+
Critical workflow rules
|
|
222
|
+
|
|
223
|
+
[MIDDLE — LOW ATTENTION]
|
|
224
|
+
Tool definitions (reference material — rarely needs full attention)
|
|
225
|
+
Desktop automation details
|
|
226
|
+
Skills system
|
|
227
|
+
|
|
228
|
+
[BOTTOM — HIGH ATTENTION]
|
|
229
|
+
Critical Rules (validation, iteration, task_complete)
|
|
230
|
+
Project awareness
|
|
231
|
+
Dynamic context (memory, git state, project files)
|
|
232
|
+
```
|
|
233
|
+
|
|
234
|
+
### 4.5 Transient vs. Persistent Application
|
|
235
|
+
|
|
236
|
+
| Scope | Mechanism | When |
|
|
237
|
+
|-------|-----------|------|
|
|
238
|
+
| **Per-session** | User sets personality with `/style concise` command | User preference |
|
|
239
|
+
| **Per-task** | Task type detection adjusts profile (code review → verbose, quick fix → concise) | Automatic |
|
|
240
|
+
| **Per-turn** | Health check prompts can inject "be more/less verbose" | Self-correction |
|
|
241
|
+
| **Persistent** | Saved to `~/.omnius/config.json` as default preference | User configuration |
|
|
242
|
+
|
|
243
|
+
### 4.6 Token Budget Implications
|
|
244
|
+
|
|
245
|
+
Per Levy et al. [Paper 2], our system prompt (~3,500 tokens) is near the reasoning degradation threshold. The personality prompt suffix should be:
|
|
246
|
+
|
|
247
|
+
- **Concise mode**: 0 extra tokens (suppress the suffix entirely)
|
|
248
|
+
- **Default mode**: 0 extra tokens (no suffix needed)
|
|
249
|
+
- **Verbose mode**: ~50 tokens max
|
|
250
|
+
- **Pedagogical mode**: ~80 tokens max
|
|
251
|
+
|
|
252
|
+
This keeps total system prompt overhead well within budget.
|
|
253
|
+
|
|
254
|
+
---
|
|
255
|
+
|
|
256
|
+
## 5. Implementation Roadmap
|
|
257
|
+
|
|
258
|
+
### Phase 1: System Prompt Position Optimization
|
|
259
|
+
- Restructure SYSTEM_PROMPT per §4.4 (critical rules at top/bottom, tools in middle)
|
|
260
|
+
- Zero new features, pure prompt restructuring
|
|
261
|
+
- Verification: eval suite pass rate maintained or improved
|
|
262
|
+
|
|
263
|
+
### Phase 2: PersonalityProfile Interface + Compiler
|
|
264
|
+
- Add PersonalityProfile type and preset profiles per §4.2
|
|
265
|
+
- Implement compilePersonalityPrompt() per §4.3
|
|
266
|
+
- Wire into AgenticRunner options
|
|
267
|
+
- Add `/style` slash command
|
|
268
|
+
|
|
269
|
+
### Phase 3: Task-Adaptive Personality
|
|
270
|
+
- Detect task type (code review, bug fix, exploration, question answering)
|
|
271
|
+
- Auto-select personality profile based on task characteristics
|
|
272
|
+
- Allow user override
|
|
273
|
+
|
|
274
|
+
### Phase 4: Self-Correcting Verbosity
|
|
275
|
+
- Health check prompts assess response style adherence
|
|
276
|
+
- Mid-task personality injection if agent drifts from target style
|
|
277
|
+
- Feedback loop with user satisfaction signal
|
|
278
|
+
|
|
279
|
+
---
|
|
280
|
+
|
|
281
|
+
## 6. References
|
|
282
|
+
|
|
283
|
+
1. Schulhoff, S., et al. (2024). "The Prompt Report: A Systematic Survey of Prompting Techniques." arXiv:2406.06608. https://arxiv.org/abs/2406.06608
|
|
284
|
+
2. Levy, M., Jacoby, A., & Goldberg, Y. (2024). "Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of LLMs." ACL 2024. https://arxiv.org/abs/2402.14848
|
|
285
|
+
3. Liu, N.F., et al. (2024). "Lost in the Middle: How Language Models Use Long Contexts." TACL 12:157–173. https://arxiv.org/abs/2307.03172
|
|
286
|
+
4. (2025). "Linear Personality Probing and Steering in LLMs: A Big Five Study." arXiv:2512.17639. https://arxiv.org/html/2512.17639v1
|
|
287
|
+
5. (2026). "SAC: A Framework for Measuring and Inducing Personality Traits in LLMs with Dynamic Intensity Control." arXiv:2506.20993. https://arxiv.org/abs/2506.20993
|
|
288
|
+
6. (2024). "Precise Length Control in Large Language Models." arXiv:2412.11937. ICLR 2026. https://arxiv.org/abs/2412.11937
|
|
289
|
+
7. (2025). "Can LLMs Track Their Output Length? A Dynamic Feedback Mechanism for Precise Length Regulation." arXiv:2601.01768. https://arxiv.org/html/2601.01768
|
|
290
|
+
8. Anthropic Engineering. (2025). "Effective Context Engineering for AI Agents." https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
|
|
291
|
+
9. (2025). "Control Illusion: The Failure of Instruction Hierarchies." arXiv:2502.15851. https://arxiv.org/pdf/2502.15851
|
|
292
|
+
10. (2025). "AgentIF: Benchmarking Instruction Following of LLMs in Agentic Scenarios." arXiv:2505.16944. https://arxiv.org/html/2505.16944v1
|
|
293
|
+
11. (2026). "Persona Prompting as a Lens on LLM Social Reasoning." arXiv:2601.20757. https://arxiv.org/abs/2601.20757
|
package/docs/rest/INDEX.md
CHANGED
|
@@ -2,6 +2,11 @@
|
|
|
2
2
|
|
|
3
3
|
This directory is the human-readable REST API reference for Omnius. It is meant to be explored incrementally by agents and humans.
|
|
4
4
|
|
|
5
|
+
For intent-first lookup, start with
|
|
6
|
+
[Discovery](./endpoints/discovery.md), `GET /v1/discovery`, or the bundled
|
|
7
|
+
[`DISCOVERY.json`](../DISCOVERY.json). Use this index after the catalog points
|
|
8
|
+
to an endpoint family.
|
|
9
|
+
|
|
5
10
|
Canonical sources:
|
|
6
11
|
|
|
7
12
|
- Runtime OpenAPI generator: `packages/cli/src/api/openapi.ts`
|
|
@@ -21,6 +26,7 @@ Use this index first, then open the smallest relevant family file:
|
|
|
21
26
|
| Auth, scopes, rate limits, API keys | `docs/rest/auth-and-scopes.md` |
|
|
22
27
|
| Errors, pagination, ETags, request IDs | `docs/rest/errors-pagination-etags.md` |
|
|
23
28
|
| OpenAPI and docs renderers | `docs/rest/openapi-source.md` |
|
|
29
|
+
| Capability and documentation discovery | `docs/rest/endpoints/discovery.md` |
|
|
24
30
|
| Chat, realtime, OpenAI-compatible inference | `docs/rest/endpoints/chat.md` |
|
|
25
31
|
| Agentic jobs and run lifecycle | `docs/rest/endpoints/run.md` |
|
|
26
32
|
| Config, endpoints, keys, profiles, projects | `docs/rest/endpoints/config.md` |
|
|
@@ -88,6 +94,7 @@ Scopes:
|
|
|
88
94
|
| Family | Representative endpoints |
|
|
89
95
|
| --- | --- |
|
|
90
96
|
| Health | `/health`, `/health/ready`, `/health/startup`, `/version`, `/metrics` |
|
|
97
|
+
| Discovery | `/v1/discovery`, `/v1/discovery/{id}` |
|
|
91
98
|
| Inference and chat | `/v1/models`, `/v1/chat/completions`, `/v1/embeddings`, `/v1/chat`, `/v1/generate`, `/api/generate`, `/v1/chat/sessions`, `/v1/chat/check-in` |
|
|
92
99
|
| AIWG | `/v1/aiwg`, `/v1/aiwg/frameworks`, `/v1/aiwg/skills`, `/v1/aiwg/use`, `/v1/aiwg/expand` |
|
|
93
100
|
| Runs | `/v1/run`, `/v1/runs`, `/v1/runs/{id}`, `/v1/todos`, `/v1/todos/{session_id}` |
|
package/docs/rest/QUICKREF.md
CHANGED
|
@@ -41,6 +41,13 @@ curl -s http://127.0.0.1:11435/health/ready
|
|
|
41
41
|
curl -s http://127.0.0.1:11435/version
|
|
42
42
|
```
|
|
43
43
|
|
|
44
|
+
## Discovery
|
|
45
|
+
|
|
46
|
+
```bash
|
|
47
|
+
curl -s 'http://127.0.0.1:11435/v1/discovery?q=bring%20your%20own%20inference'
|
|
48
|
+
curl -s http://127.0.0.1:11435/v1/discovery/tool.web-search
|
|
49
|
+
```
|
|
50
|
+
|
|
44
51
|
## Models
|
|
45
52
|
|
|
46
53
|
```bash
|
|
@@ -104,12 +111,23 @@ curl -s http://127.0.0.1:11435/v1/skills/omnius-rest-docs
|
|
|
104
111
|
|
|
105
112
|
## Tool Call
|
|
106
113
|
|
|
114
|
+
First inspect `direct_callable`:
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
curl -s http://127.0.0.1:11435/v1/tools/memory_search
|
|
118
|
+
```
|
|
119
|
+
|
|
107
120
|
```bash
|
|
108
121
|
curl -s -X POST http://127.0.0.1:11435/v1/tools/memory_search/call \
|
|
109
122
|
-H 'content-type: application/json' \
|
|
110
123
|
-d '{"args":{"query":"rest api docs"}}'
|
|
111
124
|
```
|
|
112
125
|
|
|
126
|
+
`web_search` is agent-bound. Offer it to `/v1/run`, `/v1/chat`, or
|
|
127
|
+
`/v1/chat/completions` with `agent_loop: true` and
|
|
128
|
+
`include_daemon_tools: ["read"]`; do not infer direct-call support from the
|
|
129
|
+
presence of its schema.
|
|
130
|
+
|
|
113
131
|
Deterministic bookkeeping without an agent run:
|
|
114
132
|
|
|
115
133
|
```bash
|
|
@@ -12,6 +12,7 @@
|
|
|
12
12
|
"full_reference": "docs/reference/rest-api.md",
|
|
13
13
|
"skill": ".aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md",
|
|
14
14
|
"families": [
|
|
15
|
+
{ "name": "discovery", "path": "docs/rest/endpoints/discovery.md", "triggers": ["discovery", "capabilities", "catalog", "/v1/discovery", "discover", "show"] },
|
|
15
16
|
{ "name": "chat", "path": "docs/rest/endpoints/chat.md", "triggers": ["chat", "generate", "realtime", "OpenAI-compatible", "/v1/chat", "/v1/generate"] },
|
|
16
17
|
{ "name": "run", "path": "docs/rest/endpoints/run.md", "triggers": ["run", "jobs", "todos", "scheduled"] },
|
|
17
18
|
{ "name": "config", "path": "docs/rest/endpoints/config.md", "triggers": ["config", "keys", "profiles", "projects", "endpoint"] },
|
|
@@ -24,7 +24,13 @@ key:scope:owner:rpm:tpd:max_jobs
|
|
|
24
24
|
|
|
25
25
|
Compatibility: `OMNIUS_API_KEY` and `OMNIUS_API_KEYS` are still accepted as legacy REST auth variables, but new deployments should use the `OMNIUS_REST_*` namespace.
|
|
26
26
|
|
|
27
|
-
Provider keys are separate from REST keys.
|
|
27
|
+
Provider keys are separate from REST keys. The exact upstream precedence is
|
|
28
|
+
`OMNIUS_PROVIDER_API_KEY` → `OMNIUS_MODEL_API_KEY` →
|
|
29
|
+
`OMNIUS_UPSTREAM_API_KEY` → `OMNIUS_API_KEY` → `VLLM_API_KEY` → persisted
|
|
30
|
+
endpoint configuration. `OMNIUS_API_KEY` remains a legacy fallback for both
|
|
31
|
+
provider and legacy REST auth, so avoid it in new remote deployments. Scoped
|
|
32
|
+
REST keys (`OMNIUS_REST_*` and `OMNIUS_API_KEYS`) are stripped before spawning
|
|
33
|
+
agent subprocesses and are not forwarded as model-provider credentials.
|
|
28
34
|
|
|
29
35
|
Fields:
|
|
30
36
|
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# Discovery
|
|
2
|
+
|
|
3
|
+
The discovery API exposes the same catalog shipped in `docs/DISCOVERY.json`
|
|
4
|
+
and used by `omnius discover` / `omnius show`.
|
|
5
|
+
|
|
6
|
+
| Method | Path | Purpose |
|
|
7
|
+
| --- | --- | --- |
|
|
8
|
+
| `GET` | `/v1/discovery` | Search or list catalog entries |
|
|
9
|
+
| `GET` | `/v1/discovery/{id}` | Expand one stable entry |
|
|
10
|
+
|
|
11
|
+
## Search
|
|
12
|
+
|
|
13
|
+
```bash
|
|
14
|
+
curl -s "http://127.0.0.1:11435/v1/discovery?q=web%20search&kind=tool&limit=5"
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
Query fields:
|
|
18
|
+
|
|
19
|
+
- `q`: free-text intent; omit to list entries.
|
|
20
|
+
- `kind`: one supported catalog kind.
|
|
21
|
+
- `limit`: page size.
|
|
22
|
+
- `offset`: zero-based page offset.
|
|
23
|
+
|
|
24
|
+
The response includes pagination metadata and an ETag. Use
|
|
25
|
+
`If-None-Match` when polling a long-running daemon.
|
|
26
|
+
|
|
27
|
+
## Expand
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
curl -s http://127.0.0.1:11435/v1/discovery/tool.web-search
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
An entry contains its stable ID, kind, title, summary, aliases/keywords,
|
|
34
|
+
invocation interfaces, typed references, and related entries. Discovery is
|
|
35
|
+
read-scoped and does not execute the selected capability.
|
|
36
|
+
|
|
37
|
+
Important entrypoints:
|
|
38
|
+
|
|
39
|
+
- `capability.bring-your-own-inference`
|
|
40
|
+
- `provider.anthropic`
|
|
41
|
+
- `provider.gemini`
|
|
42
|
+
- `tool.web-search`
|
|
43
|
+
- `api.tools`
|
|
44
|
+
- `operation.version-compatibility`
|
|
@@ -61,3 +61,8 @@ Usage endpoints support dashboards for requests, token counts, rate limits, and
|
|
|
61
61
|
## Ollama Pool
|
|
62
62
|
|
|
63
63
|
The pool endpoints are admin-oriented process hygiene tools. They report managed `ollama serve` processes and cleanup decisions. Cleanup accounts for stale pool state and orphan runner risks.
|
|
64
|
+
|
|
65
|
+
`GET /version` is also the compatibility bootstrap for an agent or service
|
|
66
|
+
integration. Execution clients should send `X-Omnius-Min-Version` on work
|
|
67
|
+
requests so an older daemon returns `412` before starting work. See
|
|
68
|
+
[Service Version Compatibility](../../operations/version-compatibility.md).
|
|
@@ -61,6 +61,15 @@ Profiles can also be selected with `X-Tool-Profile`. Named profiles resolve in t
|
|
|
61
61
|
|
|
62
62
|
Tool calls are gated by auth scope, tool policy, off-device policy, and profile restrictions.
|
|
63
63
|
|
|
64
|
+
Catalog presence does not imply direct invocation. Inspect
|
|
65
|
+
`direct_callable` in the live metadata. In particular, `web_search` is
|
|
66
|
+
agent-bound: discover its schema with `GET /v1/tools/web_search`, then offer it
|
|
67
|
+
through `/v1/run`, `/v1/chat`, or `/v1/chat/completions` using
|
|
68
|
+
`agent_loop: true` and `include_daemon_tools: ["read"]`. Do not assume
|
|
69
|
+
`POST /v1/tools/web_search/call` is available.
|
|
70
|
+
|
|
71
|
+
See [Tools And Web Search](../../guides/tools-and-web-search.md).
|
|
72
|
+
|
|
64
73
|
Bookkeeping tools are direct-callable and do not require a model run:
|
|
65
74
|
|
|
66
75
|
```bash
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
# Adversary System Review
|
|
2
|
+
|
|
3
|
+
**Date:** 2026-06-28
|
|
4
|
+
**Scope:** Read-only pass
|
|
5
|
+
**Files reviewed:**
|
|
6
|
+
- `packages/orchestrator/src/adversaryStream.ts` (407 lines)
|
|
7
|
+
- `packages/orchestrator/tests/adversaryStream.test.ts` (242 lines)
|
|
8
|
+
|
|
9
|
+
## Architecture
|
|
10
|
+
|
|
11
|
+
The adversary is a **generative, inference-driven auditor** running as a separate memory stream (`AdversaryStream`) adjacent to the main agent loop. It replaces the prior pure-regex adversary that only fired `false_success` when there were ZERO recent successes (missing every mixed success/failure run).
|
|
12
|
+
|
|
13
|
+
## Key Features
|
|
14
|
+
|
|
15
|
+
1. **Own persistent memory** — claims ledger (assertions, evidence state, verdicts) separate from main agent context
|
|
16
|
+
2. **Single inference per observation** — fires one backend call per high-signal claim, parses structured critique
|
|
17
|
+
3. **Classification taxonomy** — `false_failure`, `redundant_action`, `loop_detected`, plus `ok` (no intervention)
|
|
18
|
+
4. **Loop detection** — `loopSignal` tracks repeated tool calls with same target; fires when count ≥ threshold
|
|
19
|
+
5. **Success-language regex** — `SUCCESS_LANGUAGE` matches optimistic claims like "done", "verified", "live", "hydrated"
|
|
20
|
+
6. **Non-blocking** — adversary is advisory; never blocks tool calls. `adversaryRedundantSignal` is Priority 1 in critic but returns guidance, not a block.
|
|
21
|
+
|
|
22
|
+
## Integration Points
|
|
23
|
+
|
|
24
|
+
- **AgenticRunner**: manages `adversaryMode` ("backseat" | "skillcoach" | "both"), creates `AdversaryStream`, emits `debug_adversary` and `adversary_reaction` events
|
|
25
|
+
- **Critic**: `adversaryRedundantSignal` triggers WO-FIX-C path with message "The adversary recognized this exact tool call as already observed earlier."
|
|
26
|
+
- **TUI**: renders adversary buffer (last 50 entries), shows `adversary: N tracked` in status line
|
|
27
|
+
|
|
28
|
+
## Test Coverage
|
|
29
|
+
|
|
30
|
+
- `adversarySystemPrompt` encodes evidence-demanding posture (checks "started ≠ running", "exit code 0")
|
|
31
|
+
- `parseAdversaryCritique` structured output parsing
|
|
32
|
+
- `AdversaryStream` single-flight behavior, ledger persistence
|
|
33
|
+
- Loop detection with `loopSignal`
|
|
34
|
+
- `false_failure` and `redundant_action` classification
|
|
35
|
+
|
|
36
|
+
## Assessment
|
|
37
|
+
|
|
38
|
+
The system is well-structured. The generative approach (inference-driven vs regex) addresses the original blind spot. The non-blocking design is correct — adversary provides critique without preventing progress. The integration with critic's Priority 1 path gives it real influence on loop behavior.
|
|
39
|
+
|
|
40
|
+
## No Issues Found
|
|
41
|
+
|
|
42
|
+
This was a read-only review. No bugs, no improvements identified. The system is production-ready.
|