scout-ai 1.2.3 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (174) hide show
  1. checksums.yaml +4 -4
  2. data/.vimproject +138 -50
  3. data/README.md +171 -290
  4. data/Rakefile +17 -1
  5. data/VERSION +1 -1
  6. data/doc/Improvements.md +325 -0
  7. data/doc/StartHere.md +110 -0
  8. data/doc/developer/Architecture.md +126 -0
  9. data/doc/developer/Backends.md +199 -0
  10. data/doc/developer/ChatLifecycle.md +183 -0
  11. data/doc/developer/DelegationInternals.md +295 -0
  12. data/doc/developer/DesignPrinciples.md +245 -0
  13. data/doc/developer/PromptProcessing.md +292 -0
  14. data/doc/developer/Provenance.md +317 -0
  15. data/doc/user/BuildingAgents.md +345 -0
  16. data/doc/user/Cookbook.md +333 -0
  17. data/doc/user/CoreConcepts.md +181 -0
  18. data/doc/user/Delegation.md +191 -0
  19. data/doc/user/GettingStarted.md +159 -0
  20. data/doc/user/ManagingContext.md +163 -0
  21. data/doc/user/MultiAgentWorkflows.md +256 -0
  22. data/doc/user/Python.md +159 -0
  23. data/doc/user/RunningInference.md +200 -0
  24. data/doc/user/ToolCalling.md +193 -0
  25. data/doc/user/WritingChats.md +197 -0
  26. data/lib/scout/llm/agent/chat.rb +61 -11
  27. data/lib/scout/llm/agent/delegate.rb +274 -65
  28. data/lib/scout/llm/agent/iterate.rb +2 -2
  29. data/lib/scout/llm/agent/save.rb +273 -0
  30. data/lib/scout/llm/agent/workflow.rb +164 -0
  31. data/lib/scout/llm/agent.rb +86 -61
  32. data/lib/scout/llm/ask.rb +62 -17
  33. data/lib/scout/llm/backends/anthropic.rb +9 -2
  34. data/lib/scout/llm/backends/bedrock.rb +15 -3
  35. data/lib/scout/llm/backends/default.rb +183 -99
  36. data/lib/scout/llm/backends/glm.rb +58 -0
  37. data/lib/scout/llm/backends/huggingface.rb +196 -26
  38. data/lib/scout/llm/backends/ollama.rb +13 -1
  39. data/lib/scout/llm/backends/openai.rb +0 -2
  40. data/lib/scout/llm/backends/openwebui.rb +20 -13
  41. data/lib/scout/llm/backends/relay.rb +22 -22
  42. data/lib/scout/llm/backends/responses.rb +1 -1
  43. data/lib/scout/llm/chat/agent_meta.rb +264 -0
  44. data/lib/scout/llm/chat/annotation.rb +39 -10
  45. data/lib/scout/llm/chat/parse.rb +28 -6
  46. data/lib/scout/llm/chat/persist.rb +25 -0
  47. data/lib/scout/llm/chat/process/clear.rb +41 -6
  48. data/lib/scout/llm/chat/process/files.rb +21 -6
  49. data/lib/scout/llm/chat/process/meta.rb +421 -34
  50. data/lib/scout/llm/chat/process/options.rb +21 -1
  51. data/lib/scout/llm/chat/process/tools.rb +56 -15
  52. data/lib/scout/llm/chat/process.rb +4 -0
  53. data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
  54. data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
  55. data/lib/scout/llm/chat/prompt.rb +48 -0
  56. data/lib/scout/llm/chat/provenance.rb +775 -0
  57. data/lib/scout/llm/chat/tool_calls.rb +76 -0
  58. data/lib/scout/llm/chat.rb +18 -2
  59. data/lib/scout/llm/embed.rb +11 -3
  60. data/lib/scout/llm/image.rb +86 -0
  61. data/lib/scout/llm/mcp.rb +10 -2
  62. data/lib/scout/llm/rag.rb +3 -3
  63. data/lib/scout/llm/tools/call.rb +160 -11
  64. data/lib/scout/llm/tools/knowledge_base.rb +1 -1
  65. data/lib/scout/llm/tools/workflow.rb +32 -16
  66. data/lib/scout/model/python/huggingface/causal.rb +23 -5
  67. data/lib/scout/model/python/huggingface.rb +2 -1
  68. data/lib/scout-ai.rb +1 -0
  69. data/python/README.md +197 -14
  70. data/python/scout_ai/huggingface/eval.py +245 -34
  71. data/python/tests/test_huggingface_eval.py +58 -0
  72. data/research/ChatAnalyst-required-changes.md +167 -0
  73. data/research/agent-delegation-analysis.md +810 -0
  74. data/research/agent-meta-provenance-integration-plan.md +622 -0
  75. data/research/agent-workflow-analysis.md +1120 -0
  76. data/research/backends-analysis.md +836 -0
  77. data/research/chat-core-analysis.md +946 -0
  78. data/research/chatanalyst-provenance/00-baseline.md +30 -0
  79. data/research/chatanalyst-provenance/01-repo-map.md +60 -0
  80. data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
  81. data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
  82. data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
  83. data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
  84. data/research/chatanalyst-provenance/07-critic-review.md +25 -0
  85. data/research/chatanalyst-provenance/final-report.md +45 -0
  86. data/research/chatanalyst-provenance/resumption.md +37 -0
  87. data/research/coding-philosophy-analysis.md +928 -0
  88. data/research/commands-analysis.md +947 -0
  89. data/research/multi-agent-patterns-analysis.md +853 -0
  90. data/research/prompt-strategies-analysis.md +630 -0
  91. data/research/prov-verbosity-fix-notes.md +77 -0
  92. data/research/provenance-analysis.md +469 -0
  93. data/research/provenance-navigation-design.md +640 -0
  94. data/research/synthesis-report.md +487 -0
  95. data/research/tools-system-analysis.md +779 -0
  96. data/scout-ai.gemspec +100 -11
  97. data/scout_commands/agent/ask +13 -3
  98. data/scout_commands/agent/kb +2 -0
  99. data/scout_commands/llm/ask +11 -4
  100. data/scout_commands/llm/md +76 -0
  101. data/scout_commands/llm/process_queries +48 -0
  102. data/scout_commands/llm/prov +602 -0
  103. data/scout_commands/llm/word +71 -0
  104. data/scout_commands/workflow/mcp +43 -0
  105. data/share/word/reference.docx +0 -0
  106. data/test/etc/AI/mock.yaml +11 -0
  107. data/test/fixtures/backends/anthropic.json +19 -0
  108. data/test/fixtures/backends/anthropic_tool_use.json +24 -0
  109. data/test/fixtures/backends/bedrock.json +8 -0
  110. data/test/fixtures/backends/bedrock_embedding.json +3 -0
  111. data/test/fixtures/backends/bedrock_tool_use.json +17 -0
  112. data/test/fixtures/backends/ollama.json +16 -0
  113. data/test/fixtures/backends/ollama_tool_call.json +27 -0
  114. data/test/fixtures/backends/openai_chat.json +21 -0
  115. data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
  116. data/test/fixtures/backends/responses.json +33 -0
  117. data/test/fixtures/backends/responses_tool_call.json +28 -0
  118. data/test/integration/README.md +32 -0
  119. data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
  120. data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
  121. data/test/integration/scout/llm/backends/test_relay.rb +52 -0
  122. data/test/integration/scout/llm/test_infrastructure.rb +74 -0
  123. data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
  124. data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
  125. data/test/integration/scout/model/test_base.rb +91 -0
  126. data/test/scout/llm/agent/test_chat.rb +8 -2
  127. data/test/scout/llm/agent/test_save.rb +413 -0
  128. data/test/scout/llm/agent/test_workflow.rb +110 -0
  129. data/test/scout/llm/backends/test_anthropic.rb +93 -10
  130. data/test/scout/llm/backends/test_bedrock.rb +118 -2
  131. data/test/scout/llm/backends/test_huggingface.rb +137 -42
  132. data/test/scout/llm/backends/test_ollama.rb +70 -20
  133. data/test/scout/llm/backends/test_openwebui.rb +42 -40
  134. data/test/scout/llm/backends/test_relay.rb +4 -2
  135. data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
  136. data/test/scout/llm/chat/process/test_meta.rb +518 -0
  137. data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
  138. data/test/scout/llm/chat/test_agent_meta.rb +357 -0
  139. data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
  140. data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
  141. data/test/scout/llm/chat/test_parse.rb +70 -15
  142. data/test/scout/llm/chat/test_prov_cli.rb +274 -0
  143. data/test/scout/llm/chat/test_provenance.rb +240 -0
  144. data/test/scout/llm/chat/test_tool_calls.rb +38 -0
  145. data/test/scout/llm/test_agent.rb +13 -36
  146. data/test/scout/llm/test_ask.rb +75 -52
  147. data/test/scout/llm/test_chat.rb +107 -13
  148. data/test/scout/llm/test_embed.rb +48 -0
  149. data/test/scout/llm/test_rag.rb +23 -16
  150. data/test/scout/llm/test_tools.rb +12 -1
  151. data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
  152. data/test/scout/llm/tools/test_mcp.rb +5 -3
  153. data/test/scout/llm/tools/test_workflow.rb +23 -2
  154. data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
  155. data/test/scout/model/python/huggingface/test_causal.rb +9 -3
  156. data/test/scout/model/python/huggingface/test_classification.rb +11 -2
  157. data/test/scout/model/python/test_torch.rb +2 -0
  158. data/test/scout/model/python/torch/test_helpers.rb +4 -0
  159. data/test/scout/model/test_base.rb +4 -2
  160. data/test/support/availability.rb +231 -0
  161. data/test/support/fake_clients.rb +138 -0
  162. data/test/support/fixtures.rb +21 -0
  163. data/test/support/infrastructure_probes.rb +136 -0
  164. data/test/support/mock_backend.rb +215 -0
  165. data/test/test_helper.rb +32 -2
  166. metadata +99 -10
  167. data/doc/Agent.md +0 -327
  168. data/doc/Chat.md +0 -458
  169. data/doc/LLM.md +0 -340
  170. data/doc/RAG.md +0 -129
  171. data/scout_commands/documenter +0 -148
  172. data/test/scout/llm/backends/test_openai.rb +0 -192
  173. data/test/scout/llm/backends/test_responses.rb +0 -238
  174. data/test/scout/llm/test_parse.rb +0 -98
@@ -0,0 +1,199 @@
1
+ # Backends
2
+
3
+ This document describes how the LLM backend abstraction works internally:
4
+ the module composition pattern, the inference loop, error handling, and
5
+ provider differences. It is intended for framework contributors.
6
+
7
+ > For the user-facing guide on configuring inference endpoints and providers,
8
+ > see [../user/RunningInference.md](../user/RunningInference.md).
9
+ > For deep code investigation, see
10
+ > [../../research/backends-analysis.md](../../research/backends-analysis.md).
11
+
12
+ ---
13
+
14
+ ## The Backend abstraction
15
+
16
+ There is **no abstract base class**. Instead, shared logic lives in the
17
+ `LLM::Backend` module (in `lib/scout/llm/backends/default.rb`), specifically in
18
+ an inner `ClassMethods` module. Each provider is a **module** (not a class)
19
+ that composes it.
20
+
21
+ ### Composition pattern
22
+
23
+ ```ruby
24
+ module LLM
25
+ module FooMethods # provider-specific overrides
26
+ def query(...)
27
+ ...
28
+ end
29
+ end
30
+
31
+ module Foo
32
+ TAG = 'foo'
33
+ DEFAULT_MODEL = 'foo-model'
34
+
35
+ class << self
36
+ prepend FooMethods # overrides take priority
37
+ include Backend::ClassMethods # shared implementation
38
+ end
39
+ end
40
+ end
41
+ ```
42
+
43
+ Ruby's method resolution order (MRO) ensures:
44
+ - `FooMethods#query` shadows `Backend::ClassMethods#query` (via `prepend`).
45
+ - Shared methods (`ask`, `chain_tools`, `tools`, `embed`) come from `ClassMethods`.
46
+ - `self.query(...)` inside the shared `ask` method dispatches to the provider
47
+ override.
48
+
49
+ This gives every backend the full inference pipeline for free, while only
50
+ requiring it to implement the provider-specific API call and response parsing.
51
+
52
+ ---
53
+
54
+ ## The `ask` method (template method pattern)
55
+
56
+ `Backend::ClassMethods.ask` is the universal entry point for all backends. It
57
+ follows a template-method structure:
58
+
59
+ ```
60
+ ask(messages, options)
61
+ │
62
+ ├── prepare_client(options, messages) → API client setup
63
+ ├── Chat.prepare_prompt(messages, strategies) → ephemeral context management
64
+ ├── format_messages(prompt) → provider-specific formatting
65
+ ├── tools(prompt, options) → tool registry assembly
66
+ ├── query(client, formatted_prompt, tools, options) → API call
67
+ ├── process_response(...) → parse API response → message list
68
+ ├── extract_tools(response) → detect tool calls
69
+ ├── chain_tools(...) → recursive tool loop (if needed)
70
+ ├── update_meta(messages, response) → provenance annotations
71
+ └── return annotated Chat
72
+ ```
73
+
74
+ ---
75
+
76
+ ## The `chain_tools` loop
77
+
78
+ When the model emits a tool call, the backend enters a recursive loop:
79
+
80
+ ```ruby
81
+ def chain_tools(messages, output, tools, options, &block)
82
+ if output.last[:role] == 'function_call_output'
83
+ # re-call ask with the tool output appended
84
+ output + ask(messages + output, options.except(:tool_choice).merge(return_messages: true), &block)
85
+ else
86
+ output # no pending tool call — done
87
+ end
88
+ end
89
+ ```
90
+
91
+ Each iteration:
92
+ 1. Checks if the model's last message is a `function_call_output`.
93
+ 2. If so, executes the tool, appends the result, and calls `ask` again with the
94
+ growing message list.
95
+ 3. Terminates when the model's last message is a plain `assistant` message.
96
+
97
+ ### Implicit iteration limiting
98
+
99
+ There is no hard loop counter. Instead, the `shorten_tools` prompt strategy
100
+ bounds the conversation depth: tool calls beyond `MAX_TOOL_CALLS` (40) are
101
+ dropped from the prompt, and tool outputs beyond `MAX_TOOL_OUTPUTS` are
102
+ truncated or dropped. This naturally constrains how many tool-call rounds a
103
+ conversation can sustain.
104
+
105
+ ---
106
+
107
+ ## Error handling and retries
108
+
109
+ Backends wrap the API call in `begin/rescue`:
110
+
111
+ - **`BackendException`** — A custom exception class tagged with a `.chat`
112
+ accessor so callers can inspect the chat that caused the failure.
113
+ - **Retry policy** — On transient errors (rate limits, timeouts), the backend
114
+ retries with exponential backoff.
115
+ - **Agent-level exception handling** — `Agent#ask` wraps the backend call and
116
+ delegates to `@process_exception` (a user-supplied Proc) if set, which may
117
+ trigger a `retry`.
118
+
119
+ ---
120
+
121
+ ## Backend selection
122
+
123
+ `LLM.ask` selects a backend via:
124
+
125
+ 1. **Explicit `:backend` option** — `options[:backend]` selects the module.
126
+ 2. **Endpoint configuration** — Endpoints defined in
127
+ `Scout.etc.AI[<endpoint>].yaml` may specify a backend.
128
+ 3. **Hard-coded dispatch** — A `case` statement maps `:openai`, `:anthropic`,
129
+ `:responses`, `:ollama`, `:vllm` to their modules.
130
+ 4. **Dynamic loading** — Unknown backend names are resolved as a module name
131
+ (e.g., `:my_backend` → `LLM::MyBackend`), enabling third-party backends.
132
+
133
+ ---
134
+
135
+ ## Provider differences
136
+
137
+ ### OpenAI (`LLM::OpenAI`)
138
+
139
+ - **API client**: `OpenAI::Client` (ruby-openai gem).
140
+ - **Streaming**: Supports `stream_results` for token-level streaming.
141
+ - **Tool format**: `type: 'function'` with nested `function:` key.
142
+ - **Images**: Encoded as base64 `image_url` content blocks.
143
+ - **Session continuation**: Supports `previous_response_id` for
144
+ conversation-threading (Responses API).
145
+ - **Reasoning**: Extracts reasoning summaries from `o1`/`o3`-class models.
146
+
147
+ ### Anthropic (`LLM::Anthropic`)
148
+
149
+ - **API client**: HTTP client to Anthropic API.
150
+ - **Tool format**: Flat `name`, `description`, `input_schema` (no `function:`
151
+ nesting).
152
+ - **Images**: Base64 `image` content blocks with media type.
153
+ - **System messages**: Extracted from the message list and sent as a separate
154
+ parameter.
155
+ - **Reasoning**: Extracts `thinking` content blocks.
156
+
157
+ ### Ollama (`LLM::OLlama`)
158
+
159
+ - **API client**: HTTP client to local Ollama server.
160
+ - **Tool format**: Uses OpenAI-compatible format.
161
+ - **Endpoint**: Defaults to `http://localhost:11434`.
162
+ - **Embedding**: Supports embedding queries natively.
163
+
164
+ ### Responses API (`LLM::Responses`)
165
+
166
+ - **Wraps OpenAI's Responses API** — a session-oriented endpoint.
167
+ - **Session state**: Uses `previous_response_id` to maintain context
168
+ server-side, reducing token consumption.
169
+ - **Compatible with**: OpenAI's `o1`/`o3` reasoning models.
170
+
171
+ ### Bedrock (`LLM::Bedrock`)
172
+
173
+ - **Multi-provider**: Routes to different providers (Anthropic, Meta, etc.) via
174
+ the AWS Bedrock API.
175
+ - **Model naming**: Uses provider-prefixed model IDs
176
+ (e.g., `anthropic.claude-3-sonnet`).
177
+
178
+ ---
179
+
180
+ ## Key source files
181
+
182
+ | File | Responsibility |
183
+ |---|---|
184
+ | `lib/scout/llm/ask.rb` | Top-level `LLM.ask`, backend dispatch |
185
+ | `lib/scout/llm/backends/default.rb` | `Backend` module + `ClassMethods` |
186
+ | `lib/scout/llm/backends/openai.rb` | OpenAI provider |
187
+ | `lib/scout/llm/backends/anthropic.rb` | Anthropic provider |
188
+ | `lib/scout/llm/backends/ollama.rb` | Ollama provider |
189
+ | `lib/scout/llm/backends/responses.rb` | OpenAI Responses API |
190
+ | `lib/scout/llm/backends/vllm.rb` | vLLM provider |
191
+ | `lib/scout/llm/backends/bedrock.rb` | AWS Bedrock provider |
192
+
193
+ ---
194
+
195
+ ## Cross-references
196
+
197
+ - [../user/RunningInference.md](../user/RunningInference.md) — User guide for endpoints and providers.
198
+ - [PromptProcessing.md](PromptProcessing.md) — Context management integrated into the backend.
199
+ - [../../research/backends-analysis.md](../../research/backends-analysis.md) — Deep investigation.
@@ -0,0 +1,183 @@
1
+ # Chat Lifecycle
2
+
3
+ This document describes how the `Chat` abstraction works internally: its data
4
+ model, the Annotation pattern that gives it a DSL, and the compilation
5
+ pipeline that transforms chat-file text into inference-ready message arrays.
6
+ It is intended for framework contributors.
7
+
8
+ > For the user-facing chat-file format guide, see
9
+ > [../user/WritingChats.md](../user/WritingChats.md).
10
+ > For deep code investigation, see
11
+ > [../../research/chat-core-analysis.md](../../research/chat-core-analysis.md).
12
+
13
+ ---
14
+
15
+ ## The Chat-as-data philosophy
16
+
17
+ The most important design decision in Scout-AI:
18
+
19
+ > **A Chat is a plain `Array` of message `Hash`es, not an opaque object.**
20
+
21
+ The `Chat` module uses the `Annotation` pattern (from scout-essentials) to add
22
+ DSL methods to a plain Array. The underlying data structure is always directly
23
+ accessible:
24
+
25
+ ```ruby
26
+ chat = Chat.setup([])
27
+ chat.user("Hello")
28
+
29
+ chat.class # => Array
30
+ chat.first[:role] # => "user"
31
+ chat.first[:content] # => "Hello"
32
+ chat.select { |m| m[:role] == 'system' } # standard Array operations work
33
+ ```
34
+
35
+ This means chats are serializable, composable, introspectable, and cacheable,
36
+ with no lock-in.
37
+
38
+ ---
39
+
40
+ ## Message roles
41
+
42
+ Each message in a Chat is a Hash with `:role` and `:content` keys. The system
43
+ recognizes these roles:
44
+
45
+ | Role | Purpose | Visible to model? |
46
+ |---|---|---|
47
+ | `system` | System instructions | Yes |
48
+ | `user` | User message | Yes |
49
+ | `assistant` | Model response | Yes |
50
+ | `function_call` | Tool invocation request from model | Yes (as provider-specific tool_call) |
51
+ | `function_call_output` | Tool execution result. May carry a `meta` key: an Array of already-deserialized receipt field Hashes (delegated inference metadata such as `pt`/`ct`/`tt`/`inference_id`, or `job=<path>` producer references), embedded by `LLM.process_calls` when the tool returned an `LLM::Agent`. Legacy chats may instead carry a serialized `agent_meta` key; both are read, `meta` wins when both are present. See [Provenance.md](Provenance.md). | Yes (as provider-specific tool result) |
52
+ | `meta` | Provenance metadata (tokens, job references) | **No** — stripped before inference |
53
+ | `tool` | Tool definition (inline in chat) | No — extracted into tool registry |
54
+ | `introduce` | Workflow/tool introduction | No — extracted, introduces tools to the model context |
55
+ | `mcp` | MCP server declaration | No — extracted into tool registry |
56
+ | `kb` | Knowledge base declaration | No — extracted into tool registry |
57
+ | `association` | Association declaration | No — extracted |
58
+ | `option` | LLM option (model, endpoint, etc.) | No — extracted into options hash |
59
+ | `file` / `image` / `pdf` | Binary content | Processed into content blocks |
60
+
61
+ The `meta`, `tool`, `introduce`, `mcp`, `kb`, `option`, and similar roles are
62
+ **side-channel** roles: they are extracted from the message array during
63
+ compilation and do not appear in the prompt sent to the model.
64
+
65
+ ---
66
+
67
+ ## The Annotation pattern
68
+
69
+ `Chat` is not a class — it is an Annotation module:
70
+
71
+ ```ruby
72
+ module Chat
73
+ extend Annotation
74
+ # DSL methods defined here: user, system, ask, follow, option, ...
75
+ end
76
+ ```
77
+
78
+ When you call `Chat.setup(array)`, the Annotation system:
79
+
80
+ 1. Adds Chat's methods to the **singleton class** of that specific Array instance.
81
+ 2. Does **not** change the object's class (it remains `Array`).
82
+ 3. Makes the annotation **removable** via `Annotation.purge(obj)`.
83
+
84
+ This is the "annotate, don't wrap" philosophy: you get rich behavior without
85
+ sacrificing the simplicity of the underlying data type.
86
+
87
+ ---
88
+
89
+ ## The compilation pipeline
90
+
91
+ When `LLM.ask` (or `Agent#ask`) receives input, the Chat compilation pipeline
92
+ transforms it through several stages:
93
+
94
+ ```
95
+ Input (String / file / Array)
96
+ │
97
+ ▼
98
+ 1. Parse — Chat.parse: text → Array<Hash>
99
+ │ (handles role: directives, block form, indented content)
100
+ ▼
101
+ 2. Extract options — options: directives extracted into options hash
102
+ │ (model:, endpoint:, backend:, etc.)
103
+ ▼
104
+ 3. Extract tools — tool:/introduce:/mcp:/kb: roles extracted
105
+ │ into a tool registry hash
106
+ ▼
107
+ 4. Extract clear — clear: directives processed
108
+ │ (removes tool outputs from history)
109
+ ▼
110
+ 5. prepare_prompt — context strategies applied (shorten_tools)
111
+ │ EPHEMERAL: operates on a copy, never mutates stored chat
112
+ ▼
113
+ 6. format_messages — Backend translates into provider-specific format
114
+ │
115
+ ▼
116
+ API call
117
+ ```
118
+
119
+ ### Key properties
120
+
121
+ - **Side-channel extraction**: Roles like `tool`, `option`, and `meta` are
122
+ removed from the message array before the prompt is formatted. The model
123
+ never sees them.
124
+ - **Ephemeral prompt preparation**: `prepare_prompt` operates on a **local copy**
125
+ of the messages. The stored chat retains full-fidelity data. Only the
126
+ inference API call sees the shortened version.
127
+ - **Provider-specific formatting**: Each backend translates the canonical
128
+ message hashes into the format its API expects.
129
+
130
+ ---
131
+
132
+ ## Provenance annotations
133
+
134
+ After each inference, the backend inserts a `meta:` message into the chat
135
+ containing token counts and other provenance:
136
+
137
+ ```
138
+ meta: pt=1234 ct=567 tt=1801 pt_s=5000 ct_s=2000 tt_s=7000 pt_c=15000 ...
139
+ ```
140
+
141
+ These meta messages are:
142
+ - Interleaved with conversational messages in the stored chat.
143
+ - Excluded from the lineage chain (they start segments but are not provider input).
144
+ - Used by the provenance traversal system (see [Provenance.md](Provenance.md)).
145
+
146
+ ---
147
+
148
+ ## Persistence
149
+
150
+ Chats are serialized to `.chat` files — plain-text files using the chat-file
151
+ format (see [../user/WritingChats.md](../user/WritingChats.md)). The
152
+ `.chat` extension is registered as a load driver:
153
+
154
+ - **Load**: `LLM.chat(path)` or `Chat.setup(Chat.parse(File.read(path)))`.
155
+ - **Save**: `Chat.print(chat)` produces the text representation.
156
+
157
+ The format is human-readable and diffable, making it ideal for version control
158
+ and inspection.
159
+
160
+ ---
161
+
162
+ ## Key source files
163
+
164
+ | File | Responsibility |
165
+ |---|---|
166
+ | `lib/scout/llm/chat.rb` | Chat module definition, `setup`, `parse` |
167
+ | `lib/scout/llm/chat/annotation.rb` | DSL methods (user, system, ask, follow, etc.) |
168
+ | `lib/scout/llm/chat/parse.rb` | Text → Array<Hash> parser |
169
+ | `lib/scout/llm/chat/process/options.rb` | Option extraction |
170
+ | `lib/scout/llm/chat/process/tools.rb` | Tool/introduce/mcp/kb extraction |
171
+ | `lib/scout/llm/chat/process/clear.rb` | Clear directive processing |
172
+ | `lib/scout/llm/chat/process/meta.rb` | Meta messages, provenance, message_index |
173
+ | `lib/scout/llm/chat/prompt.rb` | Prompt strategies (prepare_prompt, shorten_tools) |
174
+ | `lib/scout/llm/chat/persist.rb` | .chat file load/save |
175
+
176
+ ---
177
+
178
+ ## Cross-references
179
+
180
+ - [../user/WritingChats.md](../user/WritingChats.md) — Chat-file format from the user perspective.
181
+ - [PromptProcessing.md](PromptProcessing.md) — Context management internals.
182
+ - [Provenance.md](Provenance.md) — Provenance data model.
183
+ - [../../research/chat-core-analysis.md](../../research/chat-core-analysis.md) — Deep investigation.
@@ -0,0 +1,295 @@
1
+ # Delegation Internals
2
+
3
+ This document explains how Scout-AI implements multi-agent delegation at the
4
+ code level: the `SOCIAL_INHERIT_MODES` system, the `ask` tool mechanics, the
5
+ `delegate` method, and the template-clone lifecycle. It is intended for
6
+ framework contributors.
7
+
8
+ > For the user-facing guide to delegation, see
9
+ > [../user/Delegation.md](../user/Delegation.md).
10
+ > For deep code investigation, see
11
+ > [../../research/agent-delegation-analysis.md](../../research/agent-delegation-analysis.md).
12
+
13
+ ---
14
+
15
+ ## Overview
16
+
17
+ Delegation is implemented in `lib/scout/llm/agent/delegate.rb` (323 lines). It
18
+ provides two mechanisms for one Agent to invoke another:
19
+
20
+ 1. **`socialize`** — Registers a generic `ask` tool that lets the LLM delegate
21
+ to any specialist agent by name at runtime.
22
+ 2. **`delegate`** — Registers a named `hand_off_to_<name>` tool for a specific,
23
+ pre-loaded Agent instance.
24
+
25
+ Both mechanisms build on a common infrastructure: the **template-clone
26
+ pattern**, the **socialized chat store**, and the **inheritance modes**.
27
+
28
+ ---
29
+
30
+ ## The template-clone pattern
31
+
32
+ Each specialist agent type is loaded **once** as an immutable template and stored
33
+ in `@society`:
34
+
35
+ ```
36
+ @society = {
37
+ "Worker" => Agent (template, loaded once),
38
+ "Critic" => Agent (template, loaded once)
39
+ }
40
+ ```
41
+
42
+ When a new conversation with a specialist is needed, the template is **deep
43
+ cloned** (`clone_social_agent`):
44
+
45
+ ```ruby
46
+ def clone_social_agent(template)
47
+ agent = template.clone
48
+ agent.start_chat = social_chat_copy(template.start_chat)
49
+ agent.other_options = IndiferentHash.setup(social_duplicate(template.other_options || {}))
50
+ agent.society = nil # prevent cross-contamination
51
+ agent.chats = nil # of delegation state
52
+ agent.instance_variable_set(:@current_chat, nil)
53
+ agent
54
+ end
55
+ ```
56
+
57
+ Each clone gets:
58
+ - Its own `start_chat` (deep-copied).
59
+ - Its own `other_options` (deep-copied via `social_duplicate`).
60
+ - Nilled `society` and `chats` — no accidental access to the caller's delegation state.
61
+ - A nilled `current_chat` — forces lazy re-creation.
62
+
63
+ The deep-copy (`social_duplicate`) is recursive for Hash, Array, and String,
64
+ and passes Procs by reference (since they can't be marshalled but are safe to
65
+ share).
66
+
67
+ ---
68
+
69
+ ## The socialized chat store
70
+
71
+ Live specialist instances are stored in `@chats`, keyed by `"agent_name/conversation"`:
72
+
73
+ ```
74
+ @chats = {
75
+ "Worker/default" => Agent (clone, persistent conversation),
76
+ "Worker/analysis_1" => Agent (clone, named conversation),
77
+ "Critic/default" => Agent (clone)
78
+ }
79
+ ```
80
+
81
+ Conversation keys are **scoped by agent**: `Worker/work_A` and
82
+ `Critic/work_A` are independent conversations.
83
+
84
+ ### Persisted society layout
85
+
86
+ Live conversations are also mirrored to disk, but **lazily**: nothing is
87
+ created until an agent actually saves. `Agent#save` decides the canonical
88
+ location from the save target of the *parent* agent:
89
+
90
+ - The root chat saved at `p.chat` produces a society tree rooted at
91
+ `p.chat.files/agent.society/<agent_name>/<conversation>/agent.chat`
92
+ (named agents use `<name>.society`, e.g. `worker.society`).
93
+ - A nested chat already stored at
94
+ `.../society/<a>/<c>/<file>` makes the society of *its* children the
95
+ **sibling** directory `.../society/<a>/<c>/society/...`. A nested
96
+ `agent.chat` never grows a second `.files` tree of its own. Only the
97
+ depth-0 society directory is named after the chat (`<name>.society`);
98
+ every deeper level keeps the plain `society` basename.
99
+ - `chat_task` jobs and the `scout-ai agent ask` CLI always write the agent's
100
+ own chat at `<chat_or_job_path>.files/<name>.chat`: `agent.chat` for an
101
+ unnamed/`agent` agent, `worker.chat`/`critic.chat` for named ones.
102
+ - For a persisted **chat** root that file is a full copy of the root
103
+ conversation; provenance scanning of the sidecar excludes exactly that
104
+ top-level copy, while the society conversations under
105
+ `<chat>.files/<name>.society/<agent_name>/<conversation>/agent.chat` are
106
+ included even though they are also named `agent.chat`. Job roots keep
107
+ their own top-level `<name>.chat` as a normal log node (renderers hide
108
+ it), and the `:log` relation covers exactly
109
+ `*.files/*.chat`, `*.files/*.society/**/*.chat` and the legacy
110
+ `*.files/log/**/*.chat`, so `resets/` snapshots stay out.
111
+
112
+ **Legacy layout (read-only).** Older scout-ai wrote
113
+ `<chat>.files/log/agent.chat` and the society tree under
114
+ `<chat>.files/log/society/<agent_name>/<conversation>/agent.chat`. Nothing
115
+ writes there anymore, old files are never migrated, and provenance
116
+ traversal still globs `log/**/*.chat` so chats saved by those versions
117
+ remain visible.
118
+
119
+ An agent with no live society writes only its own chat file and creates no
120
+ `.files` tree at all; parent directories are created on demand by
121
+ `Open.sensible_write`. Children get their `save_file` assigned during a
122
+ parent save, so later independent child turns keep auto-saving in place.
123
+ Saves are cycle-safe (visited paths + seen agents + a depth limit of 32) and
124
+ non-fatal: a failure is logged as a warning and the run continues.
125
+
126
+ ### `load_chat` — get-or-create
127
+
128
+ ```ruby
129
+ def load_chat(agent_name, options = {}, conversation = nil, inherit: 'tools')
130
+ key = social_chat_key(agent_name, conversation) # "Worker/work_A"
131
+ @chats[key] ||= start_social_chat(agent_name, options, inherit)
132
+ end
133
+ ```
134
+
135
+ The `inherit` parameter is only consulted **once** — when the conversation is
136
+ first created. Follow-up turns reuse the existing conversation with its
137
+ accumulated history.
138
+
139
+ ---
140
+
141
+ ## SOCIAL_INHERIT_MODES
142
+
143
+ Three modes control how much caller context flows to a specialist on first
144
+ contact:
145
+
146
+ | Mode | What is inherited | Use case |
147
+ |---|---|---|
148
+ | `none` | Nothing. Specialist starts with its own `start_chat` only. | Fully isolated sub-agent. |
149
+ | `tools` *(default)* | Tooling roles (`introduce`, `tool`, `mcp`, `kb`) from the caller's current chat. | Shared capabilities, private history. |
150
+ | `conversation` | The caller's entire current chat minus its own `start_chat` prefix. | Full context sharing for tight collaboration. |
151
+
152
+ ### Implementation: `social_inherited_context`
153
+
154
+ ```ruby
155
+ def social_inherited_context(inherit)
156
+ case inherit
157
+ when 'none'
158
+ Chat.setup([])
159
+ when 'tools'
160
+ tooling = self.current_chat.tooling
161
+ social_chat_copy(tooling)
162
+ when 'conversation'
163
+ social_caller_context
164
+ end
165
+ end
166
+ ```
167
+
168
+ ### `social_caller_context` — extracting non-start-chat messages
169
+
170
+ For `conversation` mode, the method extracts the "new" messages the caller
171
+ has added beyond its own `start_chat`. It uses a fast path (object identity
172
+ comparison) when the same Hash objects are shared, and a fallback (prefix
173
+ matching) for separately parsed Chats.
174
+
175
+ The specialist's rebuilt `start_chat` becomes:
176
+
177
+ ```
178
+ [specialist's original start_chat] + [inherited context from caller]
179
+ ```
180
+
181
+ So the specialist always gets its own system prompt first, then optionally the
182
+ caller's tools or full conversation.
183
+
184
+ ---
185
+
186
+ ## The `socialize` method
187
+
188
+ Registers a single tool named `:ask` that the LLM can invoke to delegate to
189
+ any specialist:
190
+
191
+ **Tool schema exposed to the model:**
192
+
193
+ | Parameter | Type | Required | Description |
194
+ |---|---|---|---|
195
+ | `agent` | string | Yes | Name of the specialist agent. |
196
+ | `prompt` | string | Yes | Plain-text prompt. |
197
+ | `conversation` | string | No | Named conversation ID (omit for one-shot). |
198
+ | `inherit` | enum `[none, tools, conversation]` | No (default `tools`) | Context policy for new conversations only. |
199
+
200
+ **Security boundary:** The tool block calls `ask_agent`, which uses
201
+ `agent.user(prompt)` rather than `agent.prompt(prompt)`. This is deliberate:
202
+ `prompt` parses chat-file syntax, which could allow prompt injection to inject
203
+ `tool:` or `system:` directives. The `user` method only appends a single
204
+ user-role message, making delegation safe even with untrusted LLM-generated
205
+ prompts.
206
+
207
+ **Option stripping:** Private options (`SOCIAL_PRIVATE_OPTIONS`) are stripped
208
+ before being passed to the specialist:
209
+
210
+ ```ruby
211
+ SOCIAL_PRIVATE_OPTIONS = %i[
212
+ agent client current_meta format messages no_ask_override
213
+ previous_response_id process return_messages tool_choice tools
214
+ ].freeze
215
+ ```
216
+
217
+ This prevents leaking session state, tool blocks, or message arrays from the
218
+ caller to the specialist.
219
+
220
+ ### Delegated inference receipts
221
+
222
+ When a delegation tool returns an `LLM::Agent`, `LLM.process_calls` embeds
223
+ the child agent's `meta` messages in the parent `function_call_output`
224
+ envelope under the `meta` key, as an Array of **already-deserialized** field
225
+ Hashes (one for the child's own inference metadata, one per producer
226
+ reference). These receipts are provenance evidence, not parent-chat
227
+ messages: the child's inference metadata and producer job reference are read
228
+ from the paired tool output and never injected into the parent chat. Provenance
229
+ tooling consumes them through `Chat.agent_meta_evidence` and the `:agent_job`
230
+ relation; see [Provenance.md](Provenance.md) for the extraction, precedence
231
+ (current `meta` over legacy `agent_meta`), and accounting rules.
232
+
233
+ ---
234
+
235
+ ## The `delegate` method — named hand-off
236
+
237
+ ```ruby
238
+ def delegate(agent, name, description, task_name = nil, &block)
239
+ ```
240
+
241
+ Creates a tool named `hand_off_to_#{name}` for a **specific, pre-loaded**
242
+ Agent instance. Unlike `socialize`, the agent is not chosen by the model at
243
+ call time — it is hard-coded at registration time.
244
+
245
+ | Parameter | Description |
246
+ |---|---|
247
+ | `agent` | A pre-loaded `LLM::Agent` instance. |
248
+ | `name` | Tool name suffix (e.g., `worker` → `hand_off_to_worker`). |
249
+ | `description` | Tool description for the LLM. |
250
+ | `&block` | Optional custom tool block. Defaults to: `agent.user(message); agent.chat`. |
251
+
252
+ ### Differences from `socialize`
253
+
254
+ | Aspect | `socialize` | `delegate` |
255
+ |---|---|---|
256
+ | Agent name | Model chooses at call time | Hard-coded at registration |
257
+ | Tool name | `:ask` (one tool for all agents) | `hand_off_to_#{name}` (one per agent) |
258
+ | Custom block | No (fixed block) | Yes |
259
+ | Conversation mgmt | Named conversations via `conversation` param | Single conversation, resettable via `new_conversation` |
260
+
261
+ ---
262
+
263
+ ## Legacy parameter handling
264
+
265
+ The old `chat` parameter is silently accepted for backward compatibility via
266
+ `social_tool_parameters`:
267
+
268
+ | Legacy `chat` value | Maps to `conversation` | Maps to `inherit` |
269
+ |---|---|---|
270
+ | `'current'` | `'current'` | `'conversation'` |
271
+ | `''`, `'none'`, `'false'` | `nil` (one-shot) | `'none'` |
272
+ | Any other name | That name | `'tools'` |
273
+
274
+ New code should use `conversation` and `inherit` as separate parameters.
275
+
276
+ ---
277
+
278
+ ## Key source files
279
+
280
+ | File | Responsibility |
281
+ |---|---|
282
+ | `lib/scout/llm/agent/delegate.rb` | All delegation logic |
283
+ | `lib/scout/llm/agent.rb` | `Agent` class, `ask` entry point, `load_agent` class method |
284
+ | `lib/scout/llm/agent/chat.rb` | `start_chat`, `current_chat`, Chat proxy via `method_missing` |
285
+ | `lib/scout/llm/agent/workflow.rb` | `chat_task`, `log_agent` — workflow integration |
286
+
287
+ ---
288
+
289
+ ## Cross-references
290
+
291
+ - [../user/Delegation.md](../user/Delegation.md) — User guide to delegation.
292
+ - [Provenance.md](Provenance.md) — Receipt-based provenance for delegated inference.
293
+ - [../user/MultiAgentWorkflows.md](../user/MultiAgentWorkflows.md) — Orchestration patterns.
294
+ - [../../research/agent-delegation-analysis.md](../../research/agent-delegation-analysis.md) — Deep investigation.
295
+ - [../../research/multi-agent-patterns-analysis.md](../../research/multi-agent-patterns-analysis.md) — SC26 patterns.