scout-ai 1.2.3 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.vimproject +138 -50
- data/README.md +171 -290
- data/Rakefile +17 -1
- data/VERSION +1 -1
- data/doc/Improvements.md +325 -0
- data/doc/StartHere.md +110 -0
- data/doc/developer/Architecture.md +126 -0
- data/doc/developer/Backends.md +199 -0
- data/doc/developer/ChatLifecycle.md +183 -0
- data/doc/developer/DelegationInternals.md +295 -0
- data/doc/developer/DesignPrinciples.md +245 -0
- data/doc/developer/PromptProcessing.md +292 -0
- data/doc/developer/Provenance.md +317 -0
- data/doc/user/BuildingAgents.md +345 -0
- data/doc/user/Cookbook.md +333 -0
- data/doc/user/CoreConcepts.md +181 -0
- data/doc/user/Delegation.md +191 -0
- data/doc/user/GettingStarted.md +159 -0
- data/doc/user/ManagingContext.md +163 -0
- data/doc/user/MultiAgentWorkflows.md +256 -0
- data/doc/user/Python.md +159 -0
- data/doc/user/RunningInference.md +200 -0
- data/doc/user/ToolCalling.md +193 -0
- data/doc/user/WritingChats.md +197 -0
- data/lib/scout/llm/agent/chat.rb +61 -11
- data/lib/scout/llm/agent/delegate.rb +274 -65
- data/lib/scout/llm/agent/iterate.rb +2 -2
- data/lib/scout/llm/agent/save.rb +273 -0
- data/lib/scout/llm/agent/workflow.rb +164 -0
- data/lib/scout/llm/agent.rb +86 -61
- data/lib/scout/llm/ask.rb +62 -17
- data/lib/scout/llm/backends/anthropic.rb +9 -2
- data/lib/scout/llm/backends/bedrock.rb +15 -3
- data/lib/scout/llm/backends/default.rb +183 -99
- data/lib/scout/llm/backends/glm.rb +58 -0
- data/lib/scout/llm/backends/huggingface.rb +196 -26
- data/lib/scout/llm/backends/ollama.rb +13 -1
- data/lib/scout/llm/backends/openai.rb +0 -2
- data/lib/scout/llm/backends/openwebui.rb +20 -13
- data/lib/scout/llm/backends/relay.rb +22 -22
- data/lib/scout/llm/backends/responses.rb +1 -1
- data/lib/scout/llm/chat/agent_meta.rb +264 -0
- data/lib/scout/llm/chat/annotation.rb +39 -10
- data/lib/scout/llm/chat/parse.rb +28 -6
- data/lib/scout/llm/chat/persist.rb +25 -0
- data/lib/scout/llm/chat/process/clear.rb +41 -6
- data/lib/scout/llm/chat/process/files.rb +21 -6
- data/lib/scout/llm/chat/process/meta.rb +421 -34
- data/lib/scout/llm/chat/process/options.rb +21 -1
- data/lib/scout/llm/chat/process/tools.rb +56 -15
- data/lib/scout/llm/chat/process.rb +4 -0
- data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
- data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
- data/lib/scout/llm/chat/prompt.rb +48 -0
- data/lib/scout/llm/chat/provenance.rb +775 -0
- data/lib/scout/llm/chat/tool_calls.rb +76 -0
- data/lib/scout/llm/chat.rb +18 -2
- data/lib/scout/llm/embed.rb +11 -3
- data/lib/scout/llm/image.rb +86 -0
- data/lib/scout/llm/mcp.rb +10 -2
- data/lib/scout/llm/rag.rb +3 -3
- data/lib/scout/llm/tools/call.rb +160 -11
- data/lib/scout/llm/tools/knowledge_base.rb +1 -1
- data/lib/scout/llm/tools/workflow.rb +32 -16
- data/lib/scout/model/python/huggingface/causal.rb +23 -5
- data/lib/scout/model/python/huggingface.rb +2 -1
- data/lib/scout-ai.rb +1 -0
- data/python/README.md +197 -14
- data/python/scout_ai/huggingface/eval.py +245 -34
- data/python/tests/test_huggingface_eval.py +58 -0
- data/research/ChatAnalyst-required-changes.md +167 -0
- data/research/agent-delegation-analysis.md +810 -0
- data/research/agent-meta-provenance-integration-plan.md +622 -0
- data/research/agent-workflow-analysis.md +1120 -0
- data/research/backends-analysis.md +836 -0
- data/research/chat-core-analysis.md +946 -0
- data/research/chatanalyst-provenance/00-baseline.md +30 -0
- data/research/chatanalyst-provenance/01-repo-map.md +60 -0
- data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
- data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
- data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
- data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
- data/research/chatanalyst-provenance/07-critic-review.md +25 -0
- data/research/chatanalyst-provenance/final-report.md +45 -0
- data/research/chatanalyst-provenance/resumption.md +37 -0
- data/research/coding-philosophy-analysis.md +928 -0
- data/research/commands-analysis.md +947 -0
- data/research/multi-agent-patterns-analysis.md +853 -0
- data/research/prompt-strategies-analysis.md +630 -0
- data/research/prov-verbosity-fix-notes.md +77 -0
- data/research/provenance-analysis.md +469 -0
- data/research/provenance-navigation-design.md +640 -0
- data/research/synthesis-report.md +487 -0
- data/research/tools-system-analysis.md +779 -0
- data/scout-ai.gemspec +100 -11
- data/scout_commands/agent/ask +13 -3
- data/scout_commands/agent/kb +2 -0
- data/scout_commands/llm/ask +11 -4
- data/scout_commands/llm/md +76 -0
- data/scout_commands/llm/process_queries +48 -0
- data/scout_commands/llm/prov +602 -0
- data/scout_commands/llm/word +71 -0
- data/scout_commands/workflow/mcp +43 -0
- data/share/word/reference.docx +0 -0
- data/test/etc/AI/mock.yaml +11 -0
- data/test/fixtures/backends/anthropic.json +19 -0
- data/test/fixtures/backends/anthropic_tool_use.json +24 -0
- data/test/fixtures/backends/bedrock.json +8 -0
- data/test/fixtures/backends/bedrock_embedding.json +3 -0
- data/test/fixtures/backends/bedrock_tool_use.json +17 -0
- data/test/fixtures/backends/ollama.json +16 -0
- data/test/fixtures/backends/ollama_tool_call.json +27 -0
- data/test/fixtures/backends/openai_chat.json +21 -0
- data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
- data/test/fixtures/backends/responses.json +33 -0
- data/test/fixtures/backends/responses_tool_call.json +28 -0
- data/test/integration/README.md +32 -0
- data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
- data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
- data/test/integration/scout/llm/backends/test_relay.rb +52 -0
- data/test/integration/scout/llm/test_infrastructure.rb +74 -0
- data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
- data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
- data/test/integration/scout/model/test_base.rb +91 -0
- data/test/scout/llm/agent/test_chat.rb +8 -2
- data/test/scout/llm/agent/test_save.rb +413 -0
- data/test/scout/llm/agent/test_workflow.rb +110 -0
- data/test/scout/llm/backends/test_anthropic.rb +93 -10
- data/test/scout/llm/backends/test_bedrock.rb +118 -2
- data/test/scout/llm/backends/test_huggingface.rb +137 -42
- data/test/scout/llm/backends/test_ollama.rb +70 -20
- data/test/scout/llm/backends/test_openwebui.rb +42 -40
- data/test/scout/llm/backends/test_relay.rb +4 -2
- data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
- data/test/scout/llm/chat/process/test_meta.rb +518 -0
- data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
- data/test/scout/llm/chat/test_agent_meta.rb +357 -0
- data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
- data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
- data/test/scout/llm/chat/test_parse.rb +70 -15
- data/test/scout/llm/chat/test_prov_cli.rb +274 -0
- data/test/scout/llm/chat/test_provenance.rb +240 -0
- data/test/scout/llm/chat/test_tool_calls.rb +38 -0
- data/test/scout/llm/test_agent.rb +13 -36
- data/test/scout/llm/test_ask.rb +75 -52
- data/test/scout/llm/test_chat.rb +107 -13
- data/test/scout/llm/test_embed.rb +48 -0
- data/test/scout/llm/test_rag.rb +23 -16
- data/test/scout/llm/test_tools.rb +12 -1
- data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
- data/test/scout/llm/tools/test_mcp.rb +5 -3
- data/test/scout/llm/tools/test_workflow.rb +23 -2
- data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
- data/test/scout/model/python/huggingface/test_causal.rb +9 -3
- data/test/scout/model/python/huggingface/test_classification.rb +11 -2
- data/test/scout/model/python/test_torch.rb +2 -0
- data/test/scout/model/python/torch/test_helpers.rb +4 -0
- data/test/scout/model/test_base.rb +4 -2
- data/test/support/availability.rb +231 -0
- data/test/support/fake_clients.rb +138 -0
- data/test/support/fixtures.rb +21 -0
- data/test/support/infrastructure_probes.rb +136 -0
- data/test/support/mock_backend.rb +215 -0
- data/test/test_helper.rb +32 -2
- metadata +99 -10
- data/doc/Agent.md +0 -327
- data/doc/Chat.md +0 -458
- data/doc/LLM.md +0 -340
- data/doc/RAG.md +0 -129
- data/scout_commands/documenter +0 -148
- data/test/scout/llm/backends/test_openai.rb +0 -192
- data/test/scout/llm/backends/test_responses.rb +0 -238
- data/test/scout/llm/test_parse.rb +0 -98
|
@@ -0,0 +1,779 @@
|
|
|
1
|
+
> **Disclaimer:** This is an architectural investigation, not normative
|
|
2
|
+
> documentation. It was produced during a documentation-revamp effort and may
|
|
3
|
+
> be outdated relative to the current codebase. Treat it as supporting
|
|
4
|
+
> reference material. For maintained documentation, see
|
|
5
|
+
> [../../doc/](../../doc/).
|
|
6
|
+
>
|
|
7
|
+
|
|
8
|
+
|
|
9
|
+
# 06 — Tools System
|
|
10
|
+
|
|
11
|
+
> Source files analyzed:
|
|
12
|
+
> `lib/scout/llm/tools.rb`, `lib/scout/llm/tools/call.rb`,
|
|
13
|
+
> `lib/scout/llm/tools/workflow.rb`, `lib/scout/llm/tools/knowledge_base.rb`,
|
|
14
|
+
> `lib/scout/llm/tools/mcp.rb`, `lib/scout/llm/mcp.rb`,
|
|
15
|
+
> `lib/scout/llm/backends/default.rb`, `lib/scout/llm/backends/anthropic.rb`,
|
|
16
|
+
> `lib/scout/llm/backends/openai.rb`, `lib/scout/llm/backends/ollama.rb`,
|
|
17
|
+
> `lib/scout/llm/backends/huggingface.rb`, `lib/scout/llm/backends/bedrock.rb`,
|
|
18
|
+
> `lib/scout/llm/chat/process/tools.rb`, `lib/scout/llm/ask.rb`,
|
|
19
|
+
> `lib/scout/llm/agent.rb`, `scout_commands/workflow/mcp`.
|
|
20
|
+
|
|
21
|
+
---
|
|
22
|
+
|
|
23
|
+
## 1. Tool Definition Format
|
|
24
|
+
|
|
25
|
+
### 1.1 Internal representation — the `{ name => [executor, definition] }` hash
|
|
26
|
+
|
|
27
|
+
Throughout the tools system a **tool registry** is always a `Hash` whose keys
|
|
28
|
+
are tool (function) names and whose values are a two-element array:
|
|
29
|
+
|
|
30
|
+
```ruby
|
|
31
|
+
{
|
|
32
|
+
"task_name" => [ executor_object, definition_hash ]
|
|
33
|
+
}
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
| Array slot | Meaning |
|
|
37
|
+
|---|---|
|
|
38
|
+
| `[0]` executor | The object that knows *how* to run the tool. Can be a `Workflow`, a `KnowledgeBase`, a `Proc`, a `String` (workflow name), or `nil` (fallback to block). |
|
|
39
|
+
| `[1]` definition | The JSON-schema-flavoured hash describing the tool for the LLM. When the executor itself is a `Hash` (e.g. a raw Proc definition), slot `[0]` holds the definition and slot `[1]` may be the same or different. |
|
|
40
|
+
|
|
41
|
+
When the executor in `[0]` is a `Hash`, `process_calls` treats *that hash*
|
|
42
|
+
as the definition (see `call.rb` line 36: `definition = obj if Hash === obj`).
|
|
43
|
+
|
|
44
|
+
### 1.2 The definition hash
|
|
45
|
+
|
|
46
|
+
Every definition is an `IndiferentHash` (symbol/string-indifferent access)
|
|
47
|
+
with this shape:
|
|
48
|
+
|
|
49
|
+
```ruby
|
|
50
|
+
{
|
|
51
|
+
name: "my_tool", # tool / function name
|
|
52
|
+
description: "What this tool does",
|
|
53
|
+
parameters: {
|
|
54
|
+
type: "object",
|
|
55
|
+
properties: {
|
|
56
|
+
input_a: { type: "string", description: "..." },
|
|
57
|
+
input_b: { type: "array", items: { type: "string" },
|
|
58
|
+
enum: ["x","y"], description: "..." }
|
|
59
|
+
},
|
|
60
|
+
required: ["input_a"],
|
|
61
|
+
defaults: { input_b: "x" } # optional, stripped before sending to model
|
|
62
|
+
}
|
|
63
|
+
}
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
Key points:
|
|
67
|
+
|
|
68
|
+
* `parameters` follows **JSON Schema** conventions (`type: "object"`,
|
|
69
|
+
`properties`, `required`).
|
|
70
|
+
* `defaults` is a Scout-internal extension; it is **stripped** before
|
|
71
|
+
the definition is sent to the LLM (see `default.rb` `format_tool_definitions`
|
|
72
|
+
line 229: `definition[:parameters].delete :defaults`).
|
|
73
|
+
* Some code paths store the definition **nested** under a `:function` key
|
|
74
|
+
(`{ type: 'function', function: { name:..., description:..., parameters:...} }`).
|
|
75
|
+
The `format_tool_definitions` method in each backend normalises both
|
|
76
|
+
shapes — flattening when needed.
|
|
77
|
+
|
|
78
|
+
### 1.3 How tools are registered (entry points)
|
|
79
|
+
|
|
80
|
+
Tools are gathered from three sources, all of which return the same
|
|
81
|
+
`{ name => [executor, definition] }` hash shape:
|
|
82
|
+
|
|
83
|
+
| Source | Entry point | Called from |
|
|
84
|
+
|---|---|---|
|
|
85
|
+
| **Explicit `:tools` option** | `options[:tools]` passed to `LLM.ask` or `Backend.ask` | `default.rb` `tools()` method (line 318) |
|
|
86
|
+
| **Chat message roles** (`tool`, `mcp`, `kb`, `introduce`) | `Chat.tools(messages)` / `LLM.tools(messages)` | `default.rb` `tools()` line 320 |
|
|
87
|
+
| **Agent setup** | `LLM.workflow_tools(wf)`, `LLM.knowledge_base_tool_definition(kb)` | `agent.rb` lines 89-90 |
|
|
88
|
+
|
|
89
|
+
In `Backend.ask` (default.rb line 319):
|
|
90
|
+
```ruby
|
|
91
|
+
tools = tools(formatted_prompt, options)
|
|
92
|
+
```
|
|
93
|
+
which internally does:
|
|
94
|
+
```ruby
|
|
95
|
+
def tools(messages, options)
|
|
96
|
+
tools = options.delete :tools # 1. explicit tools option
|
|
97
|
+
# normalise Array → Hash
|
|
98
|
+
tools.merge!(LLM.tools messages) # 2. tool/mcp/kb/introduce roles
|
|
99
|
+
tools.merge!(LLM.associations messages)# 3. association roles
|
|
100
|
+
tools
|
|
101
|
+
end
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
## 2. The Calling Protocol
|
|
107
|
+
|
|
108
|
+
### 2.1 Full trace: model emits tool_call → execution → result → model continues
|
|
109
|
+
|
|
110
|
+
```
|
|
111
|
+
┌─────────────────────────────────────────────────────────────┐
|
|
112
|
+
│ 1. Backend.ask() called with messages + tools │
|
|
113
|
+
│ (default.rb:475) │
|
|
114
|
+
│ │
|
|
115
|
+
│ 2. format_messages() converts internal message roles │
|
|
116
|
+
│ ('function_call', 'function_call_output') to │
|
|
117
|
+
│ provider-specific format │
|
|
118
|
+
│ (default.rb:259 format_tool_call / format_tool_output) │
|
|
119
|
+
│ │
|
|
120
|
+
│ 3. format_tool_definitions() strips :defaults, normalises │
|
|
121
|
+
│ to provider format │
|
|
122
|
+
│ │
|
|
123
|
+
│ 4. query(client, formatted_prompt, tools, options) │
|
|
124
|
+
│ → sends to LLM API │
|
|
125
|
+
│ │
|
|
126
|
+
│ 5. Model returns response containing tool_call(s) │
|
|
127
|
+
│ │
|
|
128
|
+
│ 6. process_response(messages, response, tools, options) │
|
|
129
|
+
│ a. Extracts tool_calls from response │
|
|
130
|
+
│ b. Parses each via parse_tool_call() (backend-specific) │
|
|
131
|
+
│ c. Calls LLM.process_calls(tools, tool_calls, &block) │
|
|
132
|
+
│ d. Returns output messages array │
|
|
133
|
+
│ │
|
|
134
|
+
│ 7. chain_tools() — if last message role is │
|
|
135
|
+
│ 'function_call_output', re-calls ask() so the model │
|
|
136
|
+
│ can continue with the tool results │
|
|
137
|
+
│ (default.rb:345) │
|
|
138
|
+
│ │
|
|
139
|
+
│ 8. Loop continues until model returns text (no tool_calls) │
|
|
140
|
+
└─────────────────────────────────────────────────────────────┘
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
### 2.2 `LLM.process_calls` (call.rb) — the core dispatcher
|
|
144
|
+
|
|
145
|
+
**Signature:**
|
|
146
|
+
```ruby
|
|
147
|
+
def self.process_calls(tools, calls, &block)
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
**Parameters:**
|
|
151
|
+
| Param | Type | Meaning |
|
|
152
|
+
|---|---|---|
|
|
153
|
+
| `tools` | `Hash` | The `{ name => [executor, definition] }` registry |
|
|
154
|
+
| `calls` | `Array<Hash>` | Parsed tool calls from the model response, each containing `name`, `arguments`, `id`/`call_id` |
|
|
155
|
+
| `&block` | `Proc` (optional) | Fallback executor for tools whose executor is `nil` |
|
|
156
|
+
|
|
157
|
+
**Returns:** a flat `Array` of message hashes alternating
|
|
158
|
+
`{ role: "function_call", content: ... }` and
|
|
159
|
+
`{ role: "function_call_output", content: ... }`.
|
|
160
|
+
|
|
161
|
+
**Step-by-step:**
|
|
162
|
+
|
|
163
|
+
1. **For each tool_call**, extract `tool_call_id`, `function_name`,
|
|
164
|
+
`function_arguments` via `call_id_name_and_arguments()`.
|
|
165
|
+
|
|
166
|
+
2. **Look up the tool** in the registry: `obj, definition = tools[function_name]`.
|
|
167
|
+
|
|
168
|
+
3. **Apply defaults** from `definition[:parameters][:defaults]`.
|
|
169
|
+
|
|
170
|
+
4. **Dispatch based on executor type:**
|
|
171
|
+
|
|
172
|
+
```ruby
|
|
173
|
+
case obj
|
|
174
|
+
when Proc # Lambda/Proc — call directly
|
|
175
|
+
obj.call(function_name, function_arguments)
|
|
176
|
+
when String # Workflow name string
|
|
177
|
+
# const_get or Workflow.require_workflow, then call_workflow
|
|
178
|
+
when Workflow # Workflow module
|
|
179
|
+
call_workflow(obj, function_name, function_arguments)
|
|
180
|
+
when KnowledgeBase
|
|
181
|
+
call_knowledge_base(obj, function_name, function_arguments.dup)
|
|
182
|
+
else # fallback to block
|
|
183
|
+
block.call(function_name, function_arguments)
|
|
184
|
+
end
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
5. **Handle special return types:**
|
|
188
|
+
* `Step` — queued for batch production via `Workflow.produce(jobs)`.
|
|
189
|
+
* `LLM::Agent` — batched agent chat execution (parallel via
|
|
190
|
+
`Open.traverse` with configurable `cpus`).
|
|
191
|
+
* `IO` / `TSV::Dumper` — read to string.
|
|
192
|
+
* `nil` — becomes `"success"`.
|
|
193
|
+
* `Exception` — serialised as `{ exception:, stack: }.to_json`.
|
|
194
|
+
* Other — `to_json` or `to_s`.
|
|
195
|
+
|
|
196
|
+
6. **Step/Job resolution:** After collecting all tool results, if any
|
|
197
|
+
returned a `Step`, they are produced in batch via
|
|
198
|
+
`Workflow.produce(jobs)`. Results are then loaded: `.load` if done,
|
|
199
|
+
exception JSON if errored, or force `.run` if neither.
|
|
200
|
+
|
|
201
|
+
7. **Agent resolution:** If any tool returned an `LLM::Agent`, those
|
|
202
|
+
agents are chatted in parallel (`Open.traverse`, configurable cpus
|
|
203
|
+
via `Scout::Config.get(:cpus, :agent_ask, :agents, env: 'ASK_AGENTS', default: 3)`).
|
|
204
|
+
The agent's `current_chat.follow(res)` is called to integrate the
|
|
205
|
+
response.
|
|
206
|
+
|
|
207
|
+
8. **Content truncation guard:** if the string content exceeds
|
|
208
|
+
`LLM.max_content_length` (default 100 000, configurable via
|
|
209
|
+
`Scout::Config.get(:max_content_length, :llm_tools, :tools, :llm, :ask, default: 100_000)`),
|
|
210
|
+
it is replaced with an exception JSON containing a fingerprint and
|
|
211
|
+
(if a Step was involved) the persisted file path.
|
|
212
|
+
|
|
213
|
+
9. **Output messages** are assembled as alternating pairs:
|
|
214
|
+
```ruby
|
|
215
|
+
[
|
|
216
|
+
{ role: "function_call", content: tool_call_json },
|
|
217
|
+
{ role: "function_call_output", content: { name:, content:, id: }.to_json }
|
|
218
|
+
]
|
|
219
|
+
```
|
|
220
|
+
|
|
221
|
+
### 2.3 How tool results are formatted for the model
|
|
222
|
+
|
|
223
|
+
Each backend has `format_tool_call` and `format_tool_output` methods that
|
|
224
|
+
convert the internal `function_call` / `function_call_output` roles to
|
|
225
|
+
the provider-specific wire format:
|
|
226
|
+
|
|
227
|
+
| Backend | Tool call format | Tool result format |
|
|
228
|
+
|---|---|---|
|
|
229
|
+
| **Default (OpenAI Responses API)** | `{ type: 'function_call', name:, arguments: json_string, call_id:, status: 'completed' }` | `{ type: 'function_call_output', output:, call_id: }` |
|
|
230
|
+
| **OpenAI (Chat Completions)** | `{ role: 'assistant', tool_calls: [{ type:'function', function:{name:,arguments:} }] }` | `{ role: 'tool', content:, tool_call_id: }` |
|
|
231
|
+
| **Anthropic** | `{ role: 'assistant', content: [{ type:'tool_use', id:, name:, input: }] }` | `{ role: 'user', content: [{ type:'tool_result', tool_use_id:, content: }] }` |
|
|
232
|
+
| **Ollama** | `{ role: 'assistant', tool_calls: [{ type:'function', function:{name:,arguments:} }] }` | standard tool role |
|
|
233
|
+
| **Bedrock** | provider-native | uses `LLM.tool_response` directly in-loop |
|
|
234
|
+
|
|
235
|
+
The `chain_tools` method (default.rb:345) implements the **loop continuation**:
|
|
236
|
+
if the last output message has role `function_call_output`, it re-calls
|
|
237
|
+
`ask()` so the model can see the tool results and respond again. This
|
|
238
|
+
enables multi-turn tool use within a single `LLM.ask` call.
|
|
239
|
+
|
|
240
|
+
### 2.4 `LLM.call_tools` / `LLM.tool_response` (tools.rb — legacy/simpler path)
|
|
241
|
+
|
|
242
|
+
These methods in `tools.rb` provide a **simpler, block-based** alternative
|
|
243
|
+
to `process_calls`. They are used by the Bedrock backend.
|
|
244
|
+
|
|
245
|
+
```ruby
|
|
246
|
+
def self.call_tools(tool_calls, &block)
|
|
247
|
+
# For each tool_call:
|
|
248
|
+
# 1. Calls LLM.tool_response(tool_call, &block)
|
|
249
|
+
# 2. Returns [ {role:'function_call',...}, {role:'function_call_output',...} ]
|
|
250
|
+
end
|
|
251
|
+
|
|
252
|
+
def self.tool_response(tool_call, &block)
|
|
253
|
+
# Extracts name, arguments
|
|
254
|
+
# Calls block.call(function_name, function_arguments)
|
|
255
|
+
# Returns { id:, role: "tool", content: }
|
|
256
|
+
end
|
|
257
|
+
```
|
|
258
|
+
|
|
259
|
+
### 2.5 `LLM.run_tools` (tools.rb)
|
|
260
|
+
|
|
261
|
+
```ruby
|
|
262
|
+
def self.run_tools(messages)
|
|
263
|
+
# Converts messages with role 'cmd' into role 'tool' by executing
|
|
264
|
+
# the CMD command in the content.
|
|
265
|
+
end
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
This is a **shell-command-as-tool** mechanism: messages with
|
|
269
|
+
`role: 'cmd'` have their `content` executed as a shell command and the
|
|
270
|
+
output is wrapped as a `tool` role message.
|
|
271
|
+
|
|
272
|
+
---
|
|
273
|
+
|
|
274
|
+
## 3. Workflow-as-Tools (workflow.rb)
|
|
275
|
+
|
|
276
|
+
### 3.1 Overview
|
|
277
|
+
|
|
278
|
+
Scout's `LLM.workflow_tools` introspects a `Workflow` module and exposes
|
|
279
|
+
its tasks as LLM-callable tools. Each task becomes a function whose
|
|
280
|
+
parameters are derived from the task's input declarations.
|
|
281
|
+
|
|
282
|
+
### 3.2 `LLM.task_tool_definition`
|
|
283
|
+
|
|
284
|
+
**Signature:**
|
|
285
|
+
```ruby
|
|
286
|
+
def self.task_tool_definition(workflow, task_name, inputs = nil)
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
Generates a single tool definition from a workflow task.
|
|
290
|
+
|
|
291
|
+
**Process:**
|
|
292
|
+
|
|
293
|
+
1. Retrieves `task_info` via `workflow.task_info(task_name)` — this is
|
|
294
|
+
Scout's standard task metadata containing `:inputs`,
|
|
295
|
+
`:input_types`, `:input_descriptions`, `:input_options`, `:description`.
|
|
296
|
+
|
|
297
|
+
2. **Optional `inputs` filter:** If a list of inputs is provided (either
|
|
298
|
+
as symbols or `"name=value"` strings for defaults), only those inputs
|
|
299
|
+
are exposed. Strings containing `=` set defaults:
|
|
300
|
+
```ruby
|
|
301
|
+
inputs = [:source, "threshold=0.5"]
|
|
302
|
+
# Exposes only :source, and sets default for :threshold
|
|
303
|
+
```
|
|
304
|
+
|
|
305
|
+
3. **Type mapping** via `scout_to_tool_input_type`:
|
|
306
|
+
|
|
307
|
+
| Scout input type | JSON Schema type |
|
|
308
|
+
|---|---|
|
|
309
|
+
| `:chat` | `:text` |
|
|
310
|
+
| `:text` | `:string` |
|
|
311
|
+
| `:select` | `:string` |
|
|
312
|
+
| `:path` | `:string` |
|
|
313
|
+
| `:float` | `:number` |
|
|
314
|
+
| `*_array` | `:array` |
|
|
315
|
+
| other | unchanged |
|
|
316
|
+
|
|
317
|
+
4. **Enum support:** If an input has `:select_options` in its input
|
|
318
|
+
options, these become an `"enum"` array in the schema.
|
|
319
|
+
|
|
320
|
+
5. **`return_path` injection:** For non-exec tasks, an additional
|
|
321
|
+
boolean parameter `return_path` is injected:
|
|
322
|
+
```ruby
|
|
323
|
+
properties[:return_path] = {
|
|
324
|
+
type: 'boolean',
|
|
325
|
+
description: 'Instead of the result of the job, return the path were it is persisted'
|
|
326
|
+
}
|
|
327
|
+
```
|
|
328
|
+
|
|
329
|
+
6. **Required inputs:** Only inputs explicitly marked `required: true`
|
|
330
|
+
in their input options are added to the `required` array.
|
|
331
|
+
|
|
332
|
+
7. **Returns** the definition as an `IndiferentHash`.
|
|
333
|
+
|
|
334
|
+
**Example generated definition:**
|
|
335
|
+
```ruby
|
|
336
|
+
{
|
|
337
|
+
name: :hi,
|
|
338
|
+
description: "Just say hi to someone",
|
|
339
|
+
parameters: {
|
|
340
|
+
type: "object",
|
|
341
|
+
properties: {
|
|
342
|
+
name: { type: :string, description: "Name" },
|
|
343
|
+
return_path: { type: 'boolean', description: 'Instead of...' }
|
|
344
|
+
},
|
|
345
|
+
required: ["name"]
|
|
346
|
+
}
|
|
347
|
+
}
|
|
348
|
+
```
|
|
349
|
+
|
|
350
|
+
### 3.3 `LLM.workflow_tools`
|
|
351
|
+
|
|
352
|
+
**Signature:**
|
|
353
|
+
```ruby
|
|
354
|
+
def self.workflow_tools(workflow, tasks = nil)
|
|
355
|
+
```
|
|
356
|
+
|
|
357
|
+
**Process:**
|
|
358
|
+
|
|
359
|
+
1. **If `workflow` is an Array**, recursively merges tool definitions
|
|
360
|
+
from each workflow in the array (supports multi-workflow registration).
|
|
361
|
+
|
|
362
|
+
2. **Otherwise:**
|
|
363
|
+
* Calls `Chat.allow_read_dir(workflow.directory)` to grant the chat
|
|
364
|
+
filesystem read access to the workflow source.
|
|
365
|
+
* Determines which tasks to expose:
|
|
366
|
+
* `tasks` argument if provided.
|
|
367
|
+
* `workflow.all_exports` (explicitly exported tasks) if `nil`.
|
|
368
|
+
* Falls back to `workflow.all_tasks` if no exports exist.
|
|
369
|
+
* For each task, calls `task_tool_definition` and builds the registry:
|
|
370
|
+
```ruby
|
|
371
|
+
{ task_name => [workflow, definition] }
|
|
372
|
+
```
|
|
373
|
+
|
|
374
|
+
### 3.4 `LLM.call_workflow`
|
|
375
|
+
|
|
376
|
+
**Signature:**
|
|
377
|
+
```ruby
|
|
378
|
+
def self.call_workflow(workflow, task_name, parameters = {})
|
|
379
|
+
```
|
|
380
|
+
|
|
381
|
+
Executes a workflow task when the LLM calls it.
|
|
382
|
+
|
|
383
|
+
**Process:**
|
|
384
|
+
1. Extracts special parameters: `jobname`, `return_path`, `exec_type`,
|
|
385
|
+
`allow_recursive`.
|
|
386
|
+
2. Creates a job: `workflow.job(task_name, jobname, parameters)`.
|
|
387
|
+
3. **Dispatch:**
|
|
388
|
+
* If the task is an `exec_export` or `exec_type` is `'exec'`:
|
|
389
|
+
calls `job.exec` (synchronous, in-process, no persistence).
|
|
390
|
+
* Else if `return_path` is true:
|
|
391
|
+
`job.run(true)` (async) then returns `job.path` (the file path on
|
|
392
|
+
disk where the result will be persisted).
|
|
393
|
+
* Else:
|
|
394
|
+
**Recursion guard** — raises `ScoutException` if the job is already
|
|
395
|
+
running with the current PID (prevents infinite tool-call loops).
|
|
396
|
+
Otherwise returns the `Step` (job) object for deferred production.
|
|
397
|
+
|
|
398
|
+
### 3.5 How task inputs map to tool parameters
|
|
399
|
+
|
|
400
|
+
```
|
|
401
|
+
Scout task declaration → Tool definition
|
|
402
|
+
──────────────────────────────── ──────────────────────
|
|
403
|
+
input :name, :string, "Desc" properties[:name] = {
|
|
404
|
+
type: :string,
|
|
405
|
+
description: "Desc"
|
|
406
|
+
}
|
|
407
|
+
|
|
408
|
+
input :organism, :select, "Org", properties[:organism] = {
|
|
409
|
+
select_options: ["Hsa", "Mmu"] type: :string,
|
|
410
|
+
description: "Org",
|
|
411
|
+
enum: ["Hsa", "Mmu"]
|
|
412
|
+
}
|
|
413
|
+
|
|
414
|
+
input :files, :file_array, "Files" properties[:files] = {
|
|
415
|
+
type: :array,
|
|
416
|
+
items: { type: :string },
|
|
417
|
+
description: "Files"
|
|
418
|
+
}
|
|
419
|
+
|
|
420
|
+
input :threshold, :float, properties[:threshold] = {
|
|
421
|
+
default: 0.5 type: :number,
|
|
422
|
+
description: "..."
|
|
423
|
+
}
|
|
424
|
+
# NOTE: Scout defaults are NOT
|
|
425
|
+
# automatically carried to the
|
|
426
|
+
# tool schema; only inline
|
|
427
|
+
# inputs=["threshold=0.5"] sets
|
|
428
|
+
# parameters[:defaults]
|
|
429
|
+
```
|
|
430
|
+
|
|
431
|
+
---
|
|
432
|
+
|
|
433
|
+
## 4. Knowledge Base / RAG Tools (knowledge_base.rb)
|
|
434
|
+
|
|
435
|
+
### 4.1 Overview
|
|
436
|
+
|
|
437
|
+
Scout's KnowledgeBase (an association-graph database) is exposed to the
|
|
438
|
+
LLM as a set of query tools. Each database in the KB becomes two tools:
|
|
439
|
+
one for finding associations and one for retrieving details.
|
|
440
|
+
|
|
441
|
+
### 4.2 Tools generated per database
|
|
442
|
+
|
|
443
|
+
For each database in the KB, `knowledge_base_tool_definition` generates:
|
|
444
|
+
|
|
445
|
+
#### Tool 1: Association lookup (`database_name`)
|
|
446
|
+
|
|
447
|
+
```ruby
|
|
448
|
+
# Directed database:
|
|
449
|
+
{
|
|
450
|
+
name: "gene_protein",
|
|
451
|
+
description: "Find associations for a list of entities in database gene_protein. ...
|
|
452
|
+
Returns a list in the format source~target.",
|
|
453
|
+
parameters: {
|
|
454
|
+
type: "object",
|
|
455
|
+
properties: {
|
|
456
|
+
entities: { type: "array", items: { type: :string },
|
|
457
|
+
description: 'Source entities, or targets if "reverse" is true' },
|
|
458
|
+
reverse: { type: "boolean",
|
|
459
|
+
description: 'Look for targets instead of sources, defaults to "false"' }
|
|
460
|
+
},
|
|
461
|
+
required: ["entities"]
|
|
462
|
+
}
|
|
463
|
+
}
|
|
464
|
+
|
|
465
|
+
# Undirected database:
|
|
466
|
+
# - No :reverse parameter
|
|
467
|
+
# - description mentions entity~partner format
|
|
468
|
+
```
|
|
469
|
+
|
|
470
|
+
#### Tool 2: Association details (`database_name_association_details`)
|
|
471
|
+
|
|
472
|
+
Only generated if the database has fields.
|
|
473
|
+
|
|
474
|
+
```ruby
|
|
475
|
+
# Multiple fields:
|
|
476
|
+
{
|
|
477
|
+
name: "gene_protein_association_details",
|
|
478
|
+
description: "Return details of association as a dictionary object. ...
|
|
479
|
+
The fields are: source, target, score.",
|
|
480
|
+
parameters: {
|
|
481
|
+
type: "object",
|
|
482
|
+
properties: {
|
|
483
|
+
associations: { type: "array", items: { type: :string } },
|
|
484
|
+
fields: { type: "array", items: { type: :string } }
|
|
485
|
+
},
|
|
486
|
+
required: ["associations"]
|
|
487
|
+
}
|
|
488
|
+
}
|
|
489
|
+
|
|
490
|
+
# Single field: :fields property is omitted
|
|
491
|
+
```
|
|
492
|
+
|
|
493
|
+
### 4.3 `LLM.call_knowledge_base`
|
|
494
|
+
|
|
495
|
+
**Signature:**
|
|
496
|
+
```ruby
|
|
497
|
+
def self.call_knowledge_base(knowledge_base, database, parameters = {})
|
|
498
|
+
```
|
|
499
|
+
|
|
500
|
+
**Dispatch:**
|
|
501
|
+
* If `database` ends with `_association_details`:
|
|
502
|
+
- Strips the suffix to get the real database name.
|
|
503
|
+
- Uses `knowledge_base.get_index(database)` to look up values.
|
|
504
|
+
- If `fields` given: returns `{ association => [field_values...] }`.
|
|
505
|
+
- If no fields: returns `{ association => { field => value, ... } }`.
|
|
506
|
+
|
|
507
|
+
* Otherwise (association lookup):
|
|
508
|
+
- If `reverse`: calls `knowledge_base.parents(database, entities)`.
|
|
509
|
+
- Else: calls `knowledge_base.children(database, entities)`.
|
|
510
|
+
- Returns a list of associations in `source~target` format.
|
|
511
|
+
|
|
512
|
+
### 4.4 Registration entry points
|
|
513
|
+
|
|
514
|
+
| Context | Code |
|
|
515
|
+
|---|---|
|
|
516
|
+
| **Agent** | `agent.rb:90`: `tools.merge!(LLM.knowledge_base_tool_definition(knowledge_base))` |
|
|
517
|
+
| **Chat `kb` role** | `Chat.tools()` line 200: loads KB via `KnowledgeBase.load`, generates tools |
|
|
518
|
+
| **`LLM.knowledge_base_ask`** | `ask.rb:118`: convenience method for asking questions against a KB |
|
|
519
|
+
| **`association` role** | `Chat.associations()` line 221: registers a TSV file as a KB database on-the-fly |
|
|
520
|
+
|
|
521
|
+
### 4.5 RAG integration
|
|
522
|
+
|
|
523
|
+
The KB tools are the **retrieval mechanism** for Scout's RAG pipeline.
|
|
524
|
+
The model can:
|
|
525
|
+
1. Call a database tool to find associations (e.g., gene→disease).
|
|
526
|
+
2. Call the `_association_details` tool to get field values.
|
|
527
|
+
3. Use `reverse: true` to traverse in the opposite direction.
|
|
528
|
+
|
|
529
|
+
The `association` chat role allows dynamically registering a TSV file as
|
|
530
|
+
a queryable database during a conversation:
|
|
531
|
+
```ruby
|
|
532
|
+
# In a chat message:
|
|
533
|
+
# role: association, content: "mydb /path/to/data.tsv fields=col1,col2 type=double"
|
|
534
|
+
```
|
|
535
|
+
This registers the TSV and immediately generates the corresponding tools.
|
|
536
|
+
|
|
537
|
+
---
|
|
538
|
+
|
|
539
|
+
## 5. MCP Integration (mcp.rb, tools/mcp.rb)
|
|
540
|
+
|
|
541
|
+
### 5.1 What is MCP in Scout
|
|
542
|
+
|
|
543
|
+
Scout integrates the **Model Context Protocol (MCP)** in two directions:
|
|
544
|
+
|
|
545
|
+
| Direction | File | Purpose |
|
|
546
|
+
|---|---|---|
|
|
547
|
+
| **MCP Client** (consume external tools) | `lib/scout/llm/tools/mcp.rb` | Connect to external MCP servers (HTTP or stdio) and expose their tools to Scout agents |
|
|
548
|
+
| **MCP Server** (expose Scout workflows) | `lib/scout/llm/mcp.rb` | Run a Scout workflow as an MCP server, making its tasks available to any MCP-compatible client |
|
|
549
|
+
|
|
550
|
+
Dependencies:
|
|
551
|
+
* `ruby-mcp-client` gem (the `mcp_client` require in `tools/mcp.rb`)
|
|
552
|
+
* `mcp` gem (the `mcp` require in `llm/mcp.rb`)
|
|
553
|
+
|
|
554
|
+
### 5.2 MCP Client: `LLM.mcp_tools`
|
|
555
|
+
|
|
556
|
+
**Signature:**
|
|
557
|
+
```ruby
|
|
558
|
+
def self.mcp_tools(url, options = {})
|
|
559
|
+
```
|
|
560
|
+
|
|
561
|
+
**Process:**
|
|
562
|
+
|
|
563
|
+
1. **Determine connection type:**
|
|
564
|
+
* If `url == 'stdio'`: creates a stdio-based MCP client using the
|
|
565
|
+
`command:` option.
|
|
566
|
+
* If `Open.remote?(url)`: creates an HTTP-based client. Optionally
|
|
567
|
+
adds an `Authorization: Bearer <token>` header, where the token
|
|
568
|
+
comes from `LLM.get_url_config(:key, url, :mcp)`.
|
|
569
|
+
* Otherwise: defaults based on options.
|
|
570
|
+
|
|
571
|
+
2. **Creates the client:**
|
|
572
|
+
```ruby
|
|
573
|
+
client = MCPClient.create_client(mcp_server_configs: [options.merge(type: ..., url: ...)])
|
|
574
|
+
```
|
|
575
|
+
|
|
576
|
+
3. **Lists tools** from the server: `tools = client.list_tools`.
|
|
577
|
+
|
|
578
|
+
4. **For each MCP tool**, builds a Scout tool registry entry:
|
|
579
|
+
```ruby
|
|
580
|
+
tool_definitions[name] = [block, definition]
|
|
581
|
+
```
|
|
582
|
+
Where:
|
|
583
|
+
* `definition` = `{ name:, description:, parameters: schema }` merged
|
|
584
|
+
with `{ type: 'function', function: { ... } }`.
|
|
585
|
+
* `block` = a `Proc` that calls `tool.server.call_tool(name, params)`
|
|
586
|
+
and normalises the response (extracting `content` → `text` from the
|
|
587
|
+
MCP response envelope).
|
|
588
|
+
|
|
589
|
+
5. **Response normalisation** in the block:
|
|
590
|
+
```ruby
|
|
591
|
+
res = tool.server.call_tool(name, params)
|
|
592
|
+
res = res['content'] if Hash === res && res['content']
|
|
593
|
+
res = res.first if Array === res && res.length == 1
|
|
594
|
+
res = res['content'] if Hash === res && res['content']
|
|
595
|
+
res = res['text'] if Hash === res && res['text']
|
|
596
|
+
```
|
|
597
|
+
|
|
598
|
+
### 5.3 MCP Server: `Workflow#mcp` and `Workflow#mcp_stdio`
|
|
599
|
+
|
|
600
|
+
Defined in `lib/scout/llm/mcp.rb`, these methods turn a Scout workflow
|
|
601
|
+
into an MCP server.
|
|
602
|
+
|
|
603
|
+
#### `Workflow#mcp`
|
|
604
|
+
|
|
605
|
+
```ruby
|
|
606
|
+
def mcp(*tasks)
|
|
607
|
+
# tasks defaults to all tasks if none specified
|
|
608
|
+
tools = tasks.collect do |task, inputs = nil|
|
|
609
|
+
tool_definition = LLM.task_tool_definition(self, task, inputs)
|
|
610
|
+
description = tool_definition[:description]
|
|
611
|
+
input_schema = tool_definition[:parameters].slice(:properties, :required)
|
|
612
|
+
annotations = tool_definition.slice(:title)
|
|
613
|
+
annotations[:read_only_hint] = true
|
|
614
|
+
annotations[:destructive_hint] = false
|
|
615
|
+
annotations[:idempotent_hint] = true
|
|
616
|
+
annotations[:open_world_hint] = false
|
|
617
|
+
MCP::Tool.define(name: task, description: description,
|
|
618
|
+
input_schema: input_schema,
|
|
619
|
+
annotations: annotations) do |parameters, context|
|
|
620
|
+
self.job(name, parameters).run
|
|
621
|
+
end
|
|
622
|
+
end
|
|
623
|
+
|
|
624
|
+
MCP::Server.new(name: self.name, version: "1.0.0", tools: tools)
|
|
625
|
+
end
|
|
626
|
+
```
|
|
627
|
+
|
|
628
|
+
Key details:
|
|
629
|
+
* Reuses `LLM.task_tool_definition` to generate the schema — same code
|
|
630
|
+
path as workflow-as-tools.
|
|
631
|
+
* Extracts `input_schema` as `{ properties:, required: }` (JSON Schema
|
|
632
|
+
subset, no `defaults`).
|
|
633
|
+
* Adds MCP **annotations** hinting the tools are read-only,
|
|
634
|
+
non-destructive, idempotent, and closed-world.
|
|
635
|
+
* The tool block calls `self.job(name, parameters).run`.
|
|
636
|
+
|
|
637
|
+
#### `Workflow#mcp_stdio`
|
|
638
|
+
|
|
639
|
+
```ruby
|
|
640
|
+
def mcp_stdio(*tasks)
|
|
641
|
+
server = mcp(*tasks)
|
|
642
|
+
transport = MCP::Server::Transports::StdioTransport.new(server)
|
|
643
|
+
server.transport = transport
|
|
644
|
+
transport.open
|
|
645
|
+
end
|
|
646
|
+
```
|
|
647
|
+
|
|
648
|
+
Starts the MCP server over stdio transport — the standard way to run
|
|
649
|
+
an MCP server as a subprocess.
|
|
650
|
+
|
|
651
|
+
### 5.4 CLI entry point
|
|
652
|
+
|
|
653
|
+
The `scout workflow mcp` command (`scout_commands/workflow/mcp`) runs
|
|
654
|
+
any Scout workflow as an MCP server:
|
|
655
|
+
|
|
656
|
+
```bash
|
|
657
|
+
scout workflow mcp <workflow> [<task_name>]*
|
|
658
|
+
```
|
|
659
|
+
|
|
660
|
+
If no tasks are named, exports follow the same fallback logic as
|
|
661
|
+
`workflow_tools`: explicitly exported tasks, or all tasks.
|
|
662
|
+
|
|
663
|
+
### 5.5 Chat-level MCP integration
|
|
664
|
+
|
|
665
|
+
In the chat system, MCP servers are connected via messages with
|
|
666
|
+
`role: 'mcp'` (processed in `Chat.tools()` lines 125-143):
|
|
667
|
+
|
|
668
|
+
```
|
|
669
|
+
role: mcp
|
|
670
|
+
content: https://api.example.com/mcp/ # HTTP MCP server (all tools)
|
|
671
|
+
content: https://api.example.com/mcp/ tool1 tool2 # HTTP MCP server (specific tools)
|
|
672
|
+
content: stdio my-mcp-command # stdio MCP server
|
|
673
|
+
```
|
|
674
|
+
|
|
675
|
+
---
|
|
676
|
+
|
|
677
|
+
## 6. Key Design Patterns
|
|
678
|
+
|
|
679
|
+
### 6.1 Unified tool registry shape
|
|
680
|
+
|
|
681
|
+
The single most important pattern: **every tool source produces the same
|
|
682
|
+
`{ name => [executor, definition] }` hash**, regardless of whether the
|
|
683
|
+
tool comes from a workflow task, a knowledge base database, an MCP
|
|
684
|
+
server, or a raw Proc. This means `process_calls` has one dispatch path
|
|
685
|
+
that handles all tool types.
|
|
686
|
+
|
|
687
|
+
### 6.2 Backend polymorphism via module prepend
|
|
688
|
+
|
|
689
|
+
Each backend overrides specific formatting/parsing methods:
|
|
690
|
+
```ruby
|
|
691
|
+
class << self
|
|
692
|
+
prepend OpenAIMethods # overrides
|
|
693
|
+
include Backend::ClassMethods # shared
|
|
694
|
+
end
|
|
695
|
+
```
|
|
696
|
+
|
|
697
|
+
The shared `ask()` method calls `process_response()`, which each backend
|
|
698
|
+
implements to parse its native response format and then delegates to the
|
|
699
|
+
shared `LLM.process_calls()`.
|
|
700
|
+
|
|
701
|
+
### 6.3 Deferred Step production
|
|
702
|
+
|
|
703
|
+
When a workflow tool returns a `Step` (job), it is not immediately run.
|
|
704
|
+
Instead:
|
|
705
|
+
1. All Steps are collected.
|
|
706
|
+
2. `Workflow.produce(jobs)` batch-produces them (potentially in parallel).
|
|
707
|
+
3. Results are loaded from the produced Steps.
|
|
708
|
+
|
|
709
|
+
This allows the model to call multiple workflow tools in one turn and
|
|
710
|
+
have them execute concurrently.
|
|
711
|
+
|
|
712
|
+
### 6.4 Content truncation as a safety pattern
|
|
713
|
+
|
|
714
|
+
The `max_content_length` guard (default 100K chars) prevents enormous
|
|
715
|
+
tool outputs from blowing up the model's context window. When triggered,
|
|
716
|
+
the content is replaced with a JSON error containing a `Log.fingerprint`
|
|
717
|
+
(a compact hash/summary) and, if a Step was involved, the persisted file
|
|
718
|
+
path so the model knows where the full result lives.
|
|
719
|
+
|
|
720
|
+
### 6.5 Chat-role-as-DSL for tool registration
|
|
721
|
+
|
|
722
|
+
Instead of a programmatic API, Scout uses **message roles** in the chat
|
|
723
|
+
as a declarative DSL:
|
|
724
|
+
|
|
725
|
+
| Role | Purpose |
|
|
726
|
+
|---|---|
|
|
727
|
+
| `tool` | Register a workflow task as a tool |
|
|
728
|
+
| `introduce` | Add workflow documentation to the context + register tools |
|
|
729
|
+
| `mcp` | Connect to an MCP server and expose its tools |
|
|
730
|
+
| `kb` | Load a knowledge base and expose its databases as tools |
|
|
731
|
+
| `association` | Register a TSV file as an ad-hoc KB database |
|
|
732
|
+
| `clear_tools` | Remove all tool definitions |
|
|
733
|
+
| `clear_associations` | Remove all association definitions |
|
|
734
|
+
| `cmd` | Execute a shell command (run_tools) |
|
|
735
|
+
|
|
736
|
+
These are processed by `Chat.tools()` and `Chat.associations()`, which
|
|
737
|
+
mutate the message array (consuming/removing tool-registration messages)
|
|
738
|
+
and return the accumulated tool registry.
|
|
739
|
+
|
|
740
|
+
### 6.6 Recursion guard
|
|
741
|
+
|
|
742
|
+
In `call_workflow`, a `ScoutException` is raised if a job is already
|
|
743
|
+
running with the current PID (`job.running? && job.info[:pid] == Process.pid`).
|
|
744
|
+
This prevents an agent from calling a workflow task that itself triggers
|
|
745
|
+
another agent call, which could create infinite recursion. The guard can
|
|
746
|
+
be explicitly bypassed with `allow_recursive: 'true'`.
|
|
747
|
+
|
|
748
|
+
### 6.7 Dual schema representation
|
|
749
|
+
|
|
750
|
+
Tool definitions exist in two forms:
|
|
751
|
+
1. **Internal**: flat hash `{ name:, description:, parameters: { ... } }`.
|
|
752
|
+
2. **Provider-nested**: `{ type: 'function', function: { name:, description:, parameters: { ... } } }`.
|
|
753
|
+
|
|
754
|
+
The `format_tool_definitions` method in each backend normalises between
|
|
755
|
+
these. The KB and MCP code paths produce the nested form; the workflow
|
|
756
|
+
code path produces the flat form. Both are handled transparently.
|
|
757
|
+
|
|
758
|
+
### 6.8 Method reference summary
|
|
759
|
+
|
|
760
|
+
| Method | File | Purpose |
|
|
761
|
+
|---|---|---|
|
|
762
|
+
| `LLM.call_tools` | `tools.rb` | Simple block-based tool execution (Bedrock) |
|
|
763
|
+
| `LLM.tool_response` | `tools.rb` | Execute one tool via block, format response |
|
|
764
|
+
| `LLM.run_tools` | `tools.rb` | Execute `cmd`-role messages as shell commands |
|
|
765
|
+
| `LLM.process_calls` | `tools/call.rb` | **Main dispatcher** for tool execution |
|
|
766
|
+
| `LLM.call_id_name_and_arguments` | `tools/call.rb` | Extract id/name/args from tool_call |
|
|
767
|
+
| `LLM.task_tool_definition` | `tools/workflow.rb` | Build tool definition from workflow task |
|
|
768
|
+
| `LLM.workflow_tools` | `tools/workflow.rb` | Build tool registry from workflow |
|
|
769
|
+
| `LLM.call_workflow` | `tools/workflow.rb` | Execute workflow task as tool |
|
|
770
|
+
| `LLM.scout_to_tool_input_type` | `tools/workflow.rb` | Map Scout input types to JSON Schema types |
|
|
771
|
+
| `LLM.database_tool_definition` | `tools/knowledge_base.rb` | Build association-lookup tool from KB database |
|
|
772
|
+
| `LLM.database_details_tool_definition` | `tools/knowledge_base.rb` | Build details tool from KB database |
|
|
773
|
+
| `LLM.knowledge_base_tool_definition` | `tools/knowledge_base.rb` | Build full tool registry from KB |
|
|
774
|
+
| `LLM.call_knowledge_base` | `tools/knowledge_base.rb` | Execute KB query as tool |
|
|
775
|
+
| `LLM.mcp_tools` | `tools/mcp.rb` | Connect to MCP server, build tool registry |
|
|
776
|
+
| `Workflow#mcp` | `llm/mcp.rb` | Create MCP server from workflow |
|
|
777
|
+
| `Workflow#mcp_stdio` | `llm/mcp.rb` | Start MCP server over stdio |
|
|
778
|
+
| `LLM.tools` | `chat/process/tools.rb` (via `Chat.tools`) | Process chat messages to extract tool definitions |
|
|
779
|
+
| `LLM.associations` | `chat/process/tools.rb` (via `Chat.associations`) | Process association messages |
|