scout-ai 1.2.3 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.vimproject +138 -50
- data/README.md +171 -290
- data/Rakefile +17 -1
- data/VERSION +1 -1
- data/doc/Improvements.md +325 -0
- data/doc/StartHere.md +110 -0
- data/doc/developer/Architecture.md +126 -0
- data/doc/developer/Backends.md +199 -0
- data/doc/developer/ChatLifecycle.md +183 -0
- data/doc/developer/DelegationInternals.md +295 -0
- data/doc/developer/DesignPrinciples.md +245 -0
- data/doc/developer/PromptProcessing.md +292 -0
- data/doc/developer/Provenance.md +317 -0
- data/doc/user/BuildingAgents.md +345 -0
- data/doc/user/Cookbook.md +333 -0
- data/doc/user/CoreConcepts.md +181 -0
- data/doc/user/Delegation.md +191 -0
- data/doc/user/GettingStarted.md +159 -0
- data/doc/user/ManagingContext.md +163 -0
- data/doc/user/MultiAgentWorkflows.md +256 -0
- data/doc/user/Python.md +159 -0
- data/doc/user/RunningInference.md +200 -0
- data/doc/user/ToolCalling.md +193 -0
- data/doc/user/WritingChats.md +197 -0
- data/lib/scout/llm/agent/chat.rb +61 -11
- data/lib/scout/llm/agent/delegate.rb +274 -65
- data/lib/scout/llm/agent/iterate.rb +2 -2
- data/lib/scout/llm/agent/save.rb +273 -0
- data/lib/scout/llm/agent/workflow.rb +164 -0
- data/lib/scout/llm/agent.rb +86 -61
- data/lib/scout/llm/ask.rb +62 -17
- data/lib/scout/llm/backends/anthropic.rb +9 -2
- data/lib/scout/llm/backends/bedrock.rb +15 -3
- data/lib/scout/llm/backends/default.rb +183 -99
- data/lib/scout/llm/backends/glm.rb +58 -0
- data/lib/scout/llm/backends/huggingface.rb +196 -26
- data/lib/scout/llm/backends/ollama.rb +13 -1
- data/lib/scout/llm/backends/openai.rb +0 -2
- data/lib/scout/llm/backends/openwebui.rb +20 -13
- data/lib/scout/llm/backends/relay.rb +22 -22
- data/lib/scout/llm/backends/responses.rb +1 -1
- data/lib/scout/llm/chat/agent_meta.rb +264 -0
- data/lib/scout/llm/chat/annotation.rb +39 -10
- data/lib/scout/llm/chat/parse.rb +28 -6
- data/lib/scout/llm/chat/persist.rb +25 -0
- data/lib/scout/llm/chat/process/clear.rb +41 -6
- data/lib/scout/llm/chat/process/files.rb +21 -6
- data/lib/scout/llm/chat/process/meta.rb +421 -34
- data/lib/scout/llm/chat/process/options.rb +21 -1
- data/lib/scout/llm/chat/process/tools.rb +56 -15
- data/lib/scout/llm/chat/process.rb +4 -0
- data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
- data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
- data/lib/scout/llm/chat/prompt.rb +48 -0
- data/lib/scout/llm/chat/provenance.rb +775 -0
- data/lib/scout/llm/chat/tool_calls.rb +76 -0
- data/lib/scout/llm/chat.rb +18 -2
- data/lib/scout/llm/embed.rb +11 -3
- data/lib/scout/llm/image.rb +86 -0
- data/lib/scout/llm/mcp.rb +10 -2
- data/lib/scout/llm/rag.rb +3 -3
- data/lib/scout/llm/tools/call.rb +160 -11
- data/lib/scout/llm/tools/knowledge_base.rb +1 -1
- data/lib/scout/llm/tools/workflow.rb +32 -16
- data/lib/scout/model/python/huggingface/causal.rb +23 -5
- data/lib/scout/model/python/huggingface.rb +2 -1
- data/lib/scout-ai.rb +1 -0
- data/python/README.md +197 -14
- data/python/scout_ai/huggingface/eval.py +245 -34
- data/python/tests/test_huggingface_eval.py +58 -0
- data/research/ChatAnalyst-required-changes.md +167 -0
- data/research/agent-delegation-analysis.md +810 -0
- data/research/agent-meta-provenance-integration-plan.md +622 -0
- data/research/agent-workflow-analysis.md +1120 -0
- data/research/backends-analysis.md +836 -0
- data/research/chat-core-analysis.md +946 -0
- data/research/chatanalyst-provenance/00-baseline.md +30 -0
- data/research/chatanalyst-provenance/01-repo-map.md +60 -0
- data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
- data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
- data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
- data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
- data/research/chatanalyst-provenance/07-critic-review.md +25 -0
- data/research/chatanalyst-provenance/final-report.md +45 -0
- data/research/chatanalyst-provenance/resumption.md +37 -0
- data/research/coding-philosophy-analysis.md +928 -0
- data/research/commands-analysis.md +947 -0
- data/research/multi-agent-patterns-analysis.md +853 -0
- data/research/prompt-strategies-analysis.md +630 -0
- data/research/prov-verbosity-fix-notes.md +77 -0
- data/research/provenance-analysis.md +469 -0
- data/research/provenance-navigation-design.md +640 -0
- data/research/synthesis-report.md +487 -0
- data/research/tools-system-analysis.md +779 -0
- data/scout-ai.gemspec +100 -11
- data/scout_commands/agent/ask +13 -3
- data/scout_commands/agent/kb +2 -0
- data/scout_commands/llm/ask +11 -4
- data/scout_commands/llm/md +76 -0
- data/scout_commands/llm/process_queries +48 -0
- data/scout_commands/llm/prov +602 -0
- data/scout_commands/llm/word +71 -0
- data/scout_commands/workflow/mcp +43 -0
- data/share/word/reference.docx +0 -0
- data/test/etc/AI/mock.yaml +11 -0
- data/test/fixtures/backends/anthropic.json +19 -0
- data/test/fixtures/backends/anthropic_tool_use.json +24 -0
- data/test/fixtures/backends/bedrock.json +8 -0
- data/test/fixtures/backends/bedrock_embedding.json +3 -0
- data/test/fixtures/backends/bedrock_tool_use.json +17 -0
- data/test/fixtures/backends/ollama.json +16 -0
- data/test/fixtures/backends/ollama_tool_call.json +27 -0
- data/test/fixtures/backends/openai_chat.json +21 -0
- data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
- data/test/fixtures/backends/responses.json +33 -0
- data/test/fixtures/backends/responses_tool_call.json +28 -0
- data/test/integration/README.md +32 -0
- data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
- data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
- data/test/integration/scout/llm/backends/test_relay.rb +52 -0
- data/test/integration/scout/llm/test_infrastructure.rb +74 -0
- data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
- data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
- data/test/integration/scout/model/test_base.rb +91 -0
- data/test/scout/llm/agent/test_chat.rb +8 -2
- data/test/scout/llm/agent/test_save.rb +413 -0
- data/test/scout/llm/agent/test_workflow.rb +110 -0
- data/test/scout/llm/backends/test_anthropic.rb +93 -10
- data/test/scout/llm/backends/test_bedrock.rb +118 -2
- data/test/scout/llm/backends/test_huggingface.rb +137 -42
- data/test/scout/llm/backends/test_ollama.rb +70 -20
- data/test/scout/llm/backends/test_openwebui.rb +42 -40
- data/test/scout/llm/backends/test_relay.rb +4 -2
- data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
- data/test/scout/llm/chat/process/test_meta.rb +518 -0
- data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
- data/test/scout/llm/chat/test_agent_meta.rb +357 -0
- data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
- data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
- data/test/scout/llm/chat/test_parse.rb +70 -15
- data/test/scout/llm/chat/test_prov_cli.rb +274 -0
- data/test/scout/llm/chat/test_provenance.rb +240 -0
- data/test/scout/llm/chat/test_tool_calls.rb +38 -0
- data/test/scout/llm/test_agent.rb +13 -36
- data/test/scout/llm/test_ask.rb +75 -52
- data/test/scout/llm/test_chat.rb +107 -13
- data/test/scout/llm/test_embed.rb +48 -0
- data/test/scout/llm/test_rag.rb +23 -16
- data/test/scout/llm/test_tools.rb +12 -1
- data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
- data/test/scout/llm/tools/test_mcp.rb +5 -3
- data/test/scout/llm/tools/test_workflow.rb +23 -2
- data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
- data/test/scout/model/python/huggingface/test_causal.rb +9 -3
- data/test/scout/model/python/huggingface/test_classification.rb +11 -2
- data/test/scout/model/python/test_torch.rb +2 -0
- data/test/scout/model/python/torch/test_helpers.rb +4 -0
- data/test/scout/model/test_base.rb +4 -2
- data/test/support/availability.rb +231 -0
- data/test/support/fake_clients.rb +138 -0
- data/test/support/fixtures.rb +21 -0
- data/test/support/infrastructure_probes.rb +136 -0
- data/test/support/mock_backend.rb +215 -0
- data/test/test_helper.rb +32 -2
- metadata +99 -10
- data/doc/Agent.md +0 -327
- data/doc/Chat.md +0 -458
- data/doc/LLM.md +0 -340
- data/doc/RAG.md +0 -129
- data/scout_commands/documenter +0 -148
- data/test/scout/llm/backends/test_openai.rb +0 -192
- data/test/scout/llm/backends/test_responses.rb +0 -238
- data/test/scout/llm/test_parse.rb +0 -98
|
@@ -0,0 +1,853 @@
|
|
|
1
|
+
> **Disclaimer:** This is an architectural investigation, not normative
|
|
2
|
+
> documentation. It was produced during a documentation-revamp effort and may
|
|
3
|
+
> be outdated relative to the current codebase. Treat it as supporting
|
|
4
|
+
> reference material. For maintained documentation, see
|
|
5
|
+
> [../../doc/](../../doc/).
|
|
6
|
+
>
|
|
7
|
+
|
|
8
|
+
|
|
9
|
+
# SC26 Multi-Agent Orchestration Patterns
|
|
10
|
+
|
|
11
|
+
This document extracts reusable real-world multi-agent orchestration patterns from the SC26 workflow codebase (`~/git/workflows/SC26/`). Each pattern includes the actual Ruby/Scout code, the abstractions involved, and practical guidance on when to apply it.
|
|
12
|
+
|
|
13
|
+
---
|
|
14
|
+
|
|
15
|
+
## Overview of SC26 Agent Ecosystem
|
|
16
|
+
|
|
17
|
+
### Agent roster
|
|
18
|
+
|
|
19
|
+
| Agent | Role | Has `workflow.rb`? | Has `start_chat`? | Tooling |
|
|
20
|
+
|-------|------|--------------------|--------------------|---------|
|
|
21
|
+
| **User** | Intake normalization (restating user requests) and final report synthesis. | No (uses `AgentWorkflow`) | Yes | None (`tool: false` for final report) |
|
|
22
|
+
| **Planner** | Produces candidate step-by-step plans with acceptance criteria. | No (invoked inline by Planned pipeline) | Inherits via `Agent/intro` | Same as Worker |
|
|
23
|
+
| **Searcher** | Targeted research and evidence gathering (web + docs). | No (invoked inline by Planned pipeline) | Inherits via `Agent/intro` | Web/docs search tools |
|
|
24
|
+
| **Worker** | Concrete task execution — writes code, creates artifacts, runs commands. | No | Yes | `ComputerUse`, `Skills` |
|
|
25
|
+
| **Critic** | Verifies outputs against acceptance criteria; returns PASS / NEEDS_WORK / BLOCKED. | Yes (`ask` task) | Yes | `ComputerUse` (limited) |
|
|
26
|
+
| **Manager** | Orchestrates the budgeted control loop: Search -> Edit -> Score -> Select. Delegates to all specialists via `ask`. | No (orchestrates in conversation) | Yes | `ComputerUse` (delegated), `socialize: true` |
|
|
27
|
+
| **Planned** | A complete linear pipeline: request -> search -> plan -> work -> ask. | Yes | No (pipeline defined in code) | Configurable `worker_agent` |
|
|
28
|
+
| **Branched** | Splits work into parallel sub-tasks, each handled by a separate worker. Critic aggregates. | Yes | No | Configurable `worker_agent` |
|
|
29
|
+
| **Refined** | Iterative worker-critic loop: work -> review -> repair -> repeat until PASS or BLOCKED. | Yes | No | Configurable `worker_agent` |
|
|
30
|
+
| **Analyst** | Gathers and processes data into compact, structured artifacts. | No | Yes | `ScoutCoder`, `Skills` |
|
|
31
|
+
| **InterpretData** | Two-phase pipeline: Analyst gathers artifacts -> Worker fulfills request using them. | Yes | No | Configurable `worker_agent` |
|
|
32
|
+
| **ChatAnalyst** | Inspects persisted chat sessions, agent logs, tool calls, token usage. | Yes | Yes | `ScoutCoder`, `Skills` |
|
|
33
|
+
|
|
34
|
+
### Relationship map
|
|
35
|
+
|
|
36
|
+
```
|
|
37
|
+
┌─────────────────────────────────────────┐
|
|
38
|
+
│ Manager │
|
|
39
|
+
│ (budgeted control loop orchestrator) │
|
|
40
|
+
└──┬──────┬──────┬──────┬──────┬──────┬────┘
|
|
41
|
+
│ │ │ │ │ │
|
|
42
|
+
ask │ ask │ ask │ ask │ ask │ ask
|
|
43
|
+
▼ ▼ ▼ ▼ ▼ ▼
|
|
44
|
+
User Planner Searcher Worker Critic ...
|
|
45
|
+
|
|
46
|
+
┌─────────────────────────────────────────────────────────────┐
|
|
47
|
+
│ Planned Pipeline │
|
|
48
|
+
│ request → search → plan → work → ask │
|
|
49
|
+
│ (linear dependency chain, each step is a chat_task) │
|
|
50
|
+
└─────────────────────────────────────────────────────────────┘
|
|
51
|
+
|
|
52
|
+
┌─────────────────────────────────────────────────────────────┐
|
|
53
|
+
│ Branched Pattern │
|
|
54
|
+
│ plan → spliter → [worker_1 || worker_2 || ...] → critic │
|
|
55
|
+
│ (parallel fan-out, then aggregation) │
|
|
56
|
+
└─────────────────────────────────────────────────────────────┘
|
|
57
|
+
|
|
58
|
+
┌─────────────────────────────────────────────────────────────┐
|
|
59
|
+
│ Refined Pattern │
|
|
60
|
+
│ worker → critic → (NEEDS_WORK? retry) → PASS/BLOCKED │
|
|
61
|
+
│ (iterative refinement loop using TryAgain exception) │
|
|
62
|
+
└─────────────────────────────────────────────────────────────┘
|
|
63
|
+
|
|
64
|
+
┌─────────────────────────────────────────────────────────────┐
|
|
65
|
+
│ InterpretData Pattern │
|
|
66
|
+
│ Analyst (gather artifacts) → Worker (fulfill using them) │
|
|
67
|
+
│ (data-preparation pipeline) │
|
|
68
|
+
└─────────────────────────────────────────────────────────────┘
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
---
|
|
72
|
+
|
|
73
|
+
## The Planned Pipeline Pattern
|
|
74
|
+
|
|
75
|
+
### Full structure
|
|
76
|
+
|
|
77
|
+
The `Planned` module defines a five-stage linear pipeline, each stage being a `chat_task` that depends on the previous one through Scout's `dep` mechanism:
|
|
78
|
+
|
|
79
|
+
```ruby
|
|
80
|
+
module Planned
|
|
81
|
+
extend Workflow
|
|
82
|
+
self.include_workflow AgentWorkflow
|
|
83
|
+
|
|
84
|
+
# Stage 1: Normalize the user request
|
|
85
|
+
chat_task :request do
|
|
86
|
+
agent = self.agent :User, chat: chat
|
|
87
|
+
agent.user "Restate the request from the user as self contained instructions..."
|
|
88
|
+
agent
|
|
89
|
+
end
|
|
90
|
+
|
|
91
|
+
# Stage 2: Optional search (conditionally skipped)
|
|
92
|
+
dep :request
|
|
93
|
+
chat_task :search do
|
|
94
|
+
chat = self.chat
|
|
95
|
+
chat.follow step(:request).load # context propagation
|
|
96
|
+
agent = self.agent :Searcher, chat: chat, tooling: self.tooling_intro
|
|
97
|
+
agent.user "Prepare a report that can be used as reference..."
|
|
98
|
+
agent
|
|
99
|
+
end
|
|
100
|
+
|
|
101
|
+
# Stage 3: Plan
|
|
102
|
+
dep :search
|
|
103
|
+
chat_task :plan do
|
|
104
|
+
chat = self.chat
|
|
105
|
+
chat.follow step(:request).load.last
|
|
106
|
+
chat.follow step(:search).load.last if step(:search) # optional
|
|
107
|
+
chat.message :clear_tools, true
|
|
108
|
+
agent = self.agent :Planner, chat: chat, tooling: self.tooling_intro
|
|
109
|
+
agent.user "You have been asked to fulfill a user request. Elaborate a plan"
|
|
110
|
+
agent
|
|
111
|
+
end
|
|
112
|
+
|
|
113
|
+
# Stage 4: Execute the plan
|
|
114
|
+
dep :plan
|
|
115
|
+
chat_task :work do
|
|
116
|
+
worker_agent = options[:Planned_worker_agent] || options[:worker_agent] || 'Worker'
|
|
117
|
+
chat = self.chat
|
|
118
|
+
chat.follow step(:request).load.last
|
|
119
|
+
chat.follow step(:plan).load.last
|
|
120
|
+
chat.message :clear_tools, true
|
|
121
|
+
agent = self.agent worker_agent, chat: chat, tooling: self.tooling
|
|
122
|
+
agent.user "Proceed with the plan"
|
|
123
|
+
agent
|
|
124
|
+
end
|
|
125
|
+
|
|
126
|
+
# Stage 5: Final report
|
|
127
|
+
dep :work
|
|
128
|
+
chat_task :ask do
|
|
129
|
+
chat = self.chat
|
|
130
|
+
chat.follow step(:request).load
|
|
131
|
+
chat.follow step(:plan).load
|
|
132
|
+
chat.follow step(:work).load
|
|
133
|
+
chat.message :clear_tools, true
|
|
134
|
+
agent = self.agent :User, tooling: false, chat: chat
|
|
135
|
+
agent.user "Please elaborate a final report for the User."
|
|
136
|
+
agent
|
|
137
|
+
end
|
|
138
|
+
end
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
### Task chaining via dependencies
|
|
142
|
+
|
|
143
|
+
Each `chat_task` is preceded by a `dep` declaration:
|
|
144
|
+
- `dep :request` before `:search` — search depends on request.
|
|
145
|
+
- `dep :search` before `:plan` — plan depends on search.
|
|
146
|
+
- `dep :plan` before `:work` — work depends on plan.
|
|
147
|
+
- `dep :work` before `:ask` — ask depends on work.
|
|
148
|
+
|
|
149
|
+
Scout's dependency system ensures that when you run `Planned/ask`, it automatically resolves the entire chain: `request -> search -> plan -> work -> ask`, caching each result.
|
|
150
|
+
|
|
151
|
+
### Conditional dependency skipping
|
|
152
|
+
|
|
153
|
+
The `search` step can be conditionally skipped:
|
|
154
|
+
|
|
155
|
+
```ruby
|
|
156
|
+
dep :search do |jobname, options|
|
|
157
|
+
options = LLM.options LLM.chat(options[:chat].dup)
|
|
158
|
+
if options[:use_search] == 'true'
|
|
159
|
+
{ inputs: options, jobname: jobname }
|
|
160
|
+
else
|
|
161
|
+
{ task: :request, inputs: options, jobname: jobname } # re-route to request
|
|
162
|
+
end
|
|
163
|
+
end
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
When `use_search` is not `'true'`, the dependency block returns `{task: :request, ...}`, effectively short-circuiting the search step and routing directly to `request`. Downstream steps check `if step(:search)` to conditionally follow search results.
|
|
167
|
+
|
|
168
|
+
### The `chat.follow` mechanism for context propagation
|
|
169
|
+
|
|
170
|
+
Context flows between stages via `chat.follow`:
|
|
171
|
+
|
|
172
|
+
```ruby
|
|
173
|
+
chat.follow step(:request).load # Append all messages from the request stage
|
|
174
|
+
chat.follow step(:plan).load.last # Append only the last (answer) message
|
|
175
|
+
chat.follow step(:work).load # Append all messages from the work stage
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
Key distinction:
|
|
179
|
+
- `step(:request).load` — loads the **full chat** from the request job and appends all messages.
|
|
180
|
+
- `step(:plan).load.last` — loads the chat but appends only the **final assistant message** (the plan), avoiding context bloat.
|
|
181
|
+
|
|
182
|
+
This is a critical pattern: you can choose to propagate the full conversation or just the distilled answer, depending on how much context the downstream agent needs.
|
|
183
|
+
|
|
184
|
+
### How artifacts flow between stages
|
|
185
|
+
|
|
186
|
+
- Each `chat_task` returns an `agent` object whose `.chat` (or `.answer`) is the persisted output.
|
|
187
|
+
- The next stage accesses the previous stage's output via `step(:stage_name).load`, which loads the persisted chat.
|
|
188
|
+
- `chat.follow` selectively appends messages to build the context for the next agent.
|
|
189
|
+
- `chat.message :clear_tools, true` is called before each non-worker stage to strip tool definitions from the propagated context, preventing agents from seeing irrelevant tools.
|
|
190
|
+
|
|
191
|
+
### Configurable worker agent
|
|
192
|
+
|
|
193
|
+
The pipeline supports swapping the worker:
|
|
194
|
+
|
|
195
|
+
```ruby
|
|
196
|
+
worker_agent = options[:Planned_worker_agent] || options[:worker_agent] || 'Worker'
|
|
197
|
+
```
|
|
198
|
+
|
|
199
|
+
This allows using `ScoutCoder` or any custom agent as the execution backend by passing `worker_agent=ScoutCoder`.
|
|
200
|
+
|
|
201
|
+
---
|
|
202
|
+
|
|
203
|
+
## The Manager Control Loop Pattern
|
|
204
|
+
|
|
205
|
+
### Overview
|
|
206
|
+
|
|
207
|
+
The Manager is the most sophisticated orchestration pattern. It is **not defined in a `workflow.rb`** — instead, it operates entirely through its system prompt and the `ask` tool during a live conversation. The control loop is described in prose in `start_chat` and executed by the LLM at runtime.
|
|
208
|
+
|
|
209
|
+
### Search -> Edit -> Score -> Select cycle
|
|
210
|
+
|
|
211
|
+
From the Manager system prompt:
|
|
212
|
+
|
|
213
|
+
```
|
|
214
|
+
Default control loop:
|
|
215
|
+
1. Normalize the task into explicit objectives, assumptions, missing information, success criteria, and artifacts.
|
|
216
|
+
2. Ask Planner for 2 to 4 candidate plans.
|
|
217
|
+
3. Ask Critic to score the candidate plans and recommend:
|
|
218
|
+
- a primary branch
|
|
219
|
+
- a fallback branch
|
|
220
|
+
- whether targeted search is needed before execution
|
|
221
|
+
4. Execute only one plan step at a time.
|
|
222
|
+
5. After each executed step, ask Critic to verify the result and return:
|
|
223
|
+
- status, score, missing checks, smallest next action, branch advice
|
|
224
|
+
6. If Critic returns NEEDS_WORK, prefer this order:
|
|
225
|
+
- minimal repair of the current step
|
|
226
|
+
- targeted search to unblock that step
|
|
227
|
+
- local replan of the current branch
|
|
228
|
+
- branch switch
|
|
229
|
+
7. Stop when acceptance tests pass or the budget is exhausted.
|
|
230
|
+
8. If blocked, produce the smallest set of questions for the user.
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
This is a **model-driven control loop**: the LLM decides at each turn what to do next, using the `ask` tool to delegate to specialist agents.
|
|
234
|
+
|
|
235
|
+
### Budget management (budgeted branching)
|
|
236
|
+
|
|
237
|
+
The Manager enforces explicit resource budgets:
|
|
238
|
+
|
|
239
|
+
```
|
|
240
|
+
Budget policy:
|
|
241
|
+
- Candidate plans: 2 to 4
|
|
242
|
+
- Active branches: at most 2
|
|
243
|
+
- Broad search rounds before execution: at most 1
|
|
244
|
+
- Search during execution: targeted only unless justified
|
|
245
|
+
- Repair cycles per step: at most 2
|
|
246
|
+
- Full branch switches: at most 1 unless clearly necessary
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
This prevents runaway agent loops where the system keeps trying without converging.
|
|
250
|
+
|
|
251
|
+
### Branch-specific chats
|
|
252
|
+
|
|
253
|
+
The Manager uses **named chat identifiers** to keep branch reasoning separated:
|
|
254
|
+
|
|
255
|
+
```
|
|
256
|
+
Keep branch reasoning separated with named chats when useful,
|
|
257
|
+
for example `plan_A`, `work_A`, `critic_A`.
|
|
258
|
+
```
|
|
259
|
+
|
|
260
|
+
This means the Manager can maintain parallel reasoning tracks (branch A, branch B) without cross-contaminating context. Each `ask` call specifies a `chat:` parameter to target the right conversation.
|
|
261
|
+
|
|
262
|
+
### Delegation to specialist agents
|
|
263
|
+
|
|
264
|
+
The Manager delegates through the `ask` tool. Each delegation prompt follows a structured template:
|
|
265
|
+
|
|
266
|
+
```
|
|
267
|
+
# Introduction
|
|
268
|
+
State the broader project context and why this task matters.
|
|
269
|
+
|
|
270
|
+
# Current state
|
|
271
|
+
List the relevant artifacts, assumptions, branch id, previous results, and open issues.
|
|
272
|
+
|
|
273
|
+
# Design
|
|
274
|
+
Explain how the task should be approached, including framework constraints,
|
|
275
|
+
implementation guidance, and validation expectations.
|
|
276
|
+
|
|
277
|
+
# Task
|
|
278
|
+
State exactly what the target agent must do now.
|
|
279
|
+
|
|
280
|
+
# Output required
|
|
281
|
+
Specify the expected sections, decisions, or files to return.
|
|
282
|
+
```
|
|
283
|
+
|
|
284
|
+
The Manager's `socialize: true` setting means it can see and interact with all specialist agents.
|
|
285
|
+
|
|
286
|
+
### Manager test example
|
|
287
|
+
|
|
288
|
+
The Manager test shows a real-world delegation pattern:
|
|
289
|
+
|
|
290
|
+
```
|
|
291
|
+
Use the SC26 to find out the TFs for each timepoint for each treatment and save
|
|
292
|
+
each of them in tmp/<treatment_name>.json. To avoid overpopulating the context
|
|
293
|
+
with tool calls, ask the Worker agent to process each treatment and save the
|
|
294
|
+
results separately in a new conversation. Call the Worker with ask once per treatment.
|
|
295
|
+
```
|
|
296
|
+
|
|
297
|
+
Key pattern: **batch processing via repeated `ask` calls with separate chats** — each treatment gets its own Worker conversation to avoid context overflow.
|
|
298
|
+
|
|
299
|
+
---
|
|
300
|
+
|
|
301
|
+
## The Critic Pattern
|
|
302
|
+
|
|
303
|
+
### What the Critic does
|
|
304
|
+
|
|
305
|
+
The Critic is a **verification-only agent**. It never fixes problems — it only reports them. From `start_chat`:
|
|
306
|
+
|
|
307
|
+
```
|
|
308
|
+
Your primary job is to verify results against the request and the plan.
|
|
309
|
+
|
|
310
|
+
Review principles:
|
|
311
|
+
- Be strict and evidence-based.
|
|
312
|
+
- Inspect relevant files directly when possible.
|
|
313
|
+
- Do not rely only on the summary from other agents when you can verify something.
|
|
314
|
+
- Prefer the smallest next repair.
|
|
315
|
+
- If the task is blocked by a missing technical fact, indicate that search is needed.
|
|
316
|
+
|
|
317
|
+
Do not attempt to fix a problem. Just report it.
|
|
318
|
+
```
|
|
319
|
+
|
|
320
|
+
The Critic has limited `ComputerUse` access — it can read files and run verification commands, but it is instructed not to make changes.
|
|
321
|
+
|
|
322
|
+
### PASS / NEEDS_WORK / BLOCKED decision model
|
|
323
|
+
|
|
324
|
+
The Critic returns a JSON decision:
|
|
325
|
+
|
|
326
|
+
```json
|
|
327
|
+
{
|
|
328
|
+
"name": "critic_review",
|
|
329
|
+
"type": "object",
|
|
330
|
+
"properties": {
|
|
331
|
+
"status": { "type": "string", "description": "PASS, NEEDS_WORK, or BLOCKED" },
|
|
332
|
+
"summary": { "type": "string" },
|
|
333
|
+
"issues": { "type": "array", "items": { "type": "string" }, "default": [] },
|
|
334
|
+
"next_step": { "type": "string", "default": "" },
|
|
335
|
+
"search_needed": { "type": "boolean", "default": false }
|
|
336
|
+
},
|
|
337
|
+
"required": ["status", "summary", "issues", "next_step", "search_needed"],
|
|
338
|
+
"additionalProperties": false
|
|
339
|
+
}
|
|
340
|
+
```
|
|
341
|
+
|
|
342
|
+
Three outcomes:
|
|
343
|
+
- **PASS** — Acceptance criteria met. Work is complete.
|
|
344
|
+
- **NEEDS_WORK** — Issues found but fixable. `next_step` describes the smallest repair.
|
|
345
|
+
- **BLOCKED** — Cannot proceed without external input or missing information. `search_needed` indicates whether search could help.
|
|
346
|
+
|
|
347
|
+
### Critic workflow
|
|
348
|
+
|
|
349
|
+
The `workflow.rb` for the Critic is minimal:
|
|
350
|
+
|
|
351
|
+
```ruby
|
|
352
|
+
chat_task :ask do
|
|
353
|
+
agent = self.agent :Critic, chat: chat, no_ask_override: true
|
|
354
|
+
agent.user "Evaluate the previous work."
|
|
355
|
+
response = agent.chat return_messages: true
|
|
356
|
+
begin
|
|
357
|
+
set_info :json, Chat.parse_json(response.answer)
|
|
358
|
+
rescue
|
|
359
|
+
end
|
|
360
|
+
agent
|
|
361
|
+
end
|
|
362
|
+
```
|
|
363
|
+
|
|
364
|
+
Key details:
|
|
365
|
+
- `no_ask_override: true` — prevents the Critic from being given `ask` tool capabilities (it should not delegate).
|
|
366
|
+
- `Chat.parse_json(response.answer)` — parses the Critic's JSON response and stores it as job info metadata.
|
|
367
|
+
- The Critic uses the existing `chat` context (whatever the calling workflow has accumulated).
|
|
368
|
+
|
|
369
|
+
### How scoring is consumed
|
|
370
|
+
|
|
371
|
+
Other workflows consume the Critic's JSON:
|
|
372
|
+
|
|
373
|
+
In **Refined**:
|
|
374
|
+
```ruby
|
|
375
|
+
json = critic.json
|
|
376
|
+
evaluation = IndiferentHash.setup(json)
|
|
377
|
+
case evaluation[:status]
|
|
378
|
+
when 'NEEDS_WORK' # triggers retry
|
|
379
|
+
when 'PASS' # done
|
|
380
|
+
when 'BLOCKED' # stop
|
|
381
|
+
end
|
|
382
|
+
```
|
|
383
|
+
|
|
384
|
+
In **Branched**:
|
|
385
|
+
```ruby
|
|
386
|
+
critic = self.agent :Critic, chat: chat, tooling: self.tooling
|
|
387
|
+
# ... feeds all sub-task reports ...
|
|
388
|
+
# critic reviews the aggregate
|
|
389
|
+
```
|
|
390
|
+
|
|
391
|
+
In **Manager**: the Critic's JSON is used to decide the next action in the control loop.
|
|
392
|
+
|
|
393
|
+
---
|
|
394
|
+
|
|
395
|
+
## The Branched Pattern
|
|
396
|
+
|
|
397
|
+
### What branching means
|
|
398
|
+
|
|
399
|
+
The Branched pattern **fans out** a single plan into multiple parallel sub-tasks. A "spliter" agent (an unnamed agent using the accumulated chat context) divides the work into sub-tasks, each assigned to a separate Worker in its own conversation.
|
|
400
|
+
|
|
401
|
+
### Full structure
|
|
402
|
+
|
|
403
|
+
```ruby
|
|
404
|
+
module Branched
|
|
405
|
+
extend Workflow
|
|
406
|
+
self.include_workflow AgentWorkflow
|
|
407
|
+
|
|
408
|
+
# Reuse Planned's plan task via task_alias
|
|
409
|
+
task_alias :plan, Planned, :plan
|
|
410
|
+
|
|
411
|
+
dep :plan
|
|
412
|
+
chat_task :work do
|
|
413
|
+
worker_agent = options[:Branched_worker_agent] || options[:worker_agent] || 'Worker'
|
|
414
|
+
|
|
415
|
+
chat = self.chat
|
|
416
|
+
chat.follow step(:search).load if step(:search)
|
|
417
|
+
chat.follow step(:plan).load
|
|
418
|
+
chat.message :clear_tools, true
|
|
419
|
+
|
|
420
|
+
# Phase 1: Split the work
|
|
421
|
+
spliter = self.agent nil, chat: chat # unnamed agent, uses accumulated context
|
|
422
|
+
spliter.user <<-EOF
|
|
423
|
+
Divide the job into parallel branched out sub-tasks...
|
|
424
|
+
Return them as JSON object with names and instructions...
|
|
425
|
+
EOF
|
|
426
|
+
|
|
427
|
+
plan = step(:plan).load
|
|
428
|
+
tooling = self.tooling
|
|
429
|
+
reports = {}
|
|
430
|
+
|
|
431
|
+
# Phase 2: Execute sub-tasks in parallel
|
|
432
|
+
spliter.iterate_dictionary nil, cpus: 8, bar: self.progress_bar('Branches'), into: reports do |name, instructions|
|
|
433
|
+
log name, "Start #{name}"
|
|
434
|
+
worker = self.agent worker_agent, chat: plan.dup, tooling: tooling
|
|
435
|
+
worker.user "You have been asked to produce one sub-task (#{name}):\n\n#{instructions}"
|
|
436
|
+
reports[name] = worker.chat
|
|
437
|
+
log_agent worker, "worker-#{name}"
|
|
438
|
+
log name, "Done #{name}"
|
|
439
|
+
[name, worker.chat]
|
|
440
|
+
end
|
|
441
|
+
|
|
442
|
+
log_agent spliter, 'spliter'
|
|
443
|
+
|
|
444
|
+
# Phase 3: Critic aggregates all branch reports
|
|
445
|
+
critic = self.agent :Critic, chat: chat, tooling: self.tooling
|
|
446
|
+
critic.user "The work has been complete in different subtasks. Here are the reports"
|
|
447
|
+
reports.each do |name, report|
|
|
448
|
+
critic.user "Report for sub-task #{name}:\n\n#{report}"
|
|
449
|
+
end
|
|
450
|
+
critic
|
|
451
|
+
end
|
|
452
|
+
|
|
453
|
+
dep :work
|
|
454
|
+
chat_task :ask do
|
|
455
|
+
# Final report synthesis
|
|
456
|
+
chat = self.chat
|
|
457
|
+
chat.follow step(:request).load
|
|
458
|
+
chat.follow step(:plan).load
|
|
459
|
+
chat.message :clear_tools, true
|
|
460
|
+
agent = self.agent :User, tooling: false, chat: chat
|
|
461
|
+
agent.user "The work proceeded reports and their review are here\n\n#{step(:work).load.answer}"
|
|
462
|
+
agent.user "Please elaborate a final report for the User."
|
|
463
|
+
agent
|
|
464
|
+
end
|
|
465
|
+
end
|
|
466
|
+
```
|
|
467
|
+
|
|
468
|
+
### Key abstractions
|
|
469
|
+
|
|
470
|
+
1. **`task_alias`** — reuses another workflow's task:
|
|
471
|
+
```ruby
|
|
472
|
+
task_alias :plan, Planned, :plan
|
|
473
|
+
```
|
|
474
|
+
This makes `Branched/plan` an alias for `Planned/plan`, inheriting its dependency chain (`request` -> `search` -> `plan`).
|
|
475
|
+
|
|
476
|
+
2. **Unnamed spliter agent** — `self.agent nil, chat: chat` creates an agent without a specific system prompt. It uses whatever context has been accumulated in the chat. Its job is purely structural: divide work into sub-tasks.
|
|
477
|
+
|
|
478
|
+
3. **Parallel execution via `iterate_dictionary`**:
|
|
479
|
+
```ruby
|
|
480
|
+
spliter.iterate_dictionary nil, cpus: 8, bar: self.progress_bar('Branches'), into: reports do |name, instructions|
|
|
481
|
+
```
|
|
482
|
+
This iterates over the JSON dictionary returned by the spliter, spawning a Worker per entry with up to 8 concurrent executions. Each Worker gets:
|
|
483
|
+
- A **fresh chat** from the plan (`chat: plan.dup`) — not the full accumulated context.
|
|
484
|
+
- The sub-task-specific instructions.
|
|
485
|
+
|
|
486
|
+
4. **Branch isolation** — each worker gets `plan.dup` (a copy of the plan chat), ensuring branches don't interfere with each other. This is critical for parallel safety.
|
|
487
|
+
|
|
488
|
+
5. **Aggregation by Critic** — after all branches complete, a single Critic agent receives all reports sequentially:
|
|
489
|
+
```ruby
|
|
490
|
+
reports.each do |name, report|
|
|
491
|
+
critic.user "Report for sub-task #{name}:\n\n#{report}"
|
|
492
|
+
end
|
|
493
|
+
```
|
|
494
|
+
|
|
495
|
+
### When to use
|
|
496
|
+
|
|
497
|
+
- When a task can be decomposed into independent sub-tasks (e.g., processing each treatment in a dataset separately).
|
|
498
|
+
- When parallelism provides speedup (up to `cpus: 8` concurrent workers).
|
|
499
|
+
- When sub-tasks share the same plan but operate on different data.
|
|
500
|
+
|
|
501
|
+
---
|
|
502
|
+
|
|
503
|
+
## The Refined Pattern
|
|
504
|
+
|
|
505
|
+
### What refinement means
|
|
506
|
+
|
|
507
|
+
The Refined pattern is an **iterative worker-critic loop**. Unlike branching (which fans out horizontally), refinement iterates vertically: the same worker attempts the task, the critic reviews, and if the result is not good enough, the worker tries again with the critic's feedback.
|
|
508
|
+
|
|
509
|
+
### Full structure
|
|
510
|
+
|
|
511
|
+
```ruby
|
|
512
|
+
module Refined
|
|
513
|
+
extend Workflow
|
|
514
|
+
self.include_workflow AgentWorkflow
|
|
515
|
+
|
|
516
|
+
chat_task :ask do
|
|
517
|
+
worker_agent = options[:Refined_worker_agent] || options[:worker_agent] || 'Worker'
|
|
518
|
+
|
|
519
|
+
chat = self.chat
|
|
520
|
+
worker = self.agent worker_agent, chat: chat, tooling: self.tooling
|
|
521
|
+
critic = self.agent :Critic, chat: chat
|
|
522
|
+
|
|
523
|
+
round = 1
|
|
524
|
+
begin
|
|
525
|
+
# Worker executes
|
|
526
|
+
worker.user "Execute the work you have been assigned and write a report for the Critic agent."
|
|
527
|
+
report = worker.chat
|
|
528
|
+
|
|
529
|
+
# Critic evaluates
|
|
530
|
+
critic.user "Below is the workers report"
|
|
531
|
+
critic.user report
|
|
532
|
+
critic.user "Make an evaluation in JSON."
|
|
533
|
+
json = critic.json
|
|
534
|
+
evaluation = IndiferentHash.setup(json)
|
|
535
|
+
|
|
536
|
+
log_agent worker, "worker-round-#{round}"
|
|
537
|
+
log_agent critic, "critic-round-#{round}"
|
|
538
|
+
|
|
539
|
+
case evaluation[:status]
|
|
540
|
+
when 'NEEDS_WORK'
|
|
541
|
+
round += 1
|
|
542
|
+
# Feed evaluation back to worker
|
|
543
|
+
worker.user "The Critic has done this evaluation and proposed more work."
|
|
544
|
+
worker.user evaluation.to_json
|
|
545
|
+
worker.message :clear_tools, true
|
|
546
|
+
critic.message :clear_tools, true
|
|
547
|
+
raise TryAgain # exception-based retry
|
|
548
|
+
|
|
549
|
+
when 'PASS'
|
|
550
|
+
# Done — fall through
|
|
551
|
+
|
|
552
|
+
when 'BLOCKED'
|
|
553
|
+
# Cannot proceed — fall through
|
|
554
|
+
end
|
|
555
|
+
rescue
|
|
556
|
+
retry if TryAgain === $! # catch TryAgain and retry the begin block
|
|
557
|
+
raise $!
|
|
558
|
+
end
|
|
559
|
+
|
|
560
|
+
{ role: :assistant, content: evaluation.to_json }
|
|
561
|
+
end
|
|
562
|
+
end
|
|
563
|
+
```
|
|
564
|
+
|
|
565
|
+
### How it differs from branching
|
|
566
|
+
|
|
567
|
+
| Aspect | Refined | Branched |
|
|
568
|
+
|--------|---------|----------|
|
|
569
|
+
| **Direction** | Vertical (iterative deepening) | Horizontal (parallel fan-out) |
|
|
570
|
+
| **Worker count** | 1 worker, multiple rounds | N workers, 1 round each |
|
|
571
|
+
| **Critic role** | After each round, triggers retry | Once, after all branches complete |
|
|
572
|
+
| **Loop control** | `TryAgain` exception + `retry` | `iterate_dictionary` with `cpus` |
|
|
573
|
+
| **Context** | Shared chat accumulates across rounds | Each worker gets `plan.dup` (isolated) |
|
|
574
|
+
| **Termination** | Critic says PASS or BLOCKED | All sub-tasks complete |
|
|
575
|
+
|
|
576
|
+
### The `TryAgain` exception pattern
|
|
577
|
+
|
|
578
|
+
This is a Scout-ism for controlled retries:
|
|
579
|
+
|
|
580
|
+
```ruby
|
|
581
|
+
raise TryAgain # inside the begin block
|
|
582
|
+
# ...
|
|
583
|
+
rescue
|
|
584
|
+
retry if TryAgain === $! # catches it and re-executes the begin block
|
|
585
|
+
raise $! # re-raises any other exception
|
|
586
|
+
```
|
|
587
|
+
|
|
588
|
+
The `retry` keyword re-executes the entire `begin...end` block. State (like `round`) is preserved because it's declared outside the `begin`. This gives the worker a fresh attempt while keeping the accumulated chat context (worker and critic share the same `chat`).
|
|
589
|
+
|
|
590
|
+
### Clear tools between rounds
|
|
591
|
+
|
|
592
|
+
```ruby
|
|
593
|
+
worker.message :clear_tools, true
|
|
594
|
+
critic.message :clear_tools, true
|
|
595
|
+
```
|
|
596
|
+
|
|
597
|
+
After each NEEDS_WORK cycle, tools are cleared from the chat context. This prevents tool definitions from accumulating and bloating the context window across iterations.
|
|
598
|
+
|
|
599
|
+
---
|
|
600
|
+
|
|
601
|
+
## The InterpretData Pattern
|
|
602
|
+
|
|
603
|
+
### Overview
|
|
604
|
+
|
|
605
|
+
InterpretData is a **data-preparation pipeline** that bridges an Analyst agent (who reduces large data to compact artifacts) with a Worker agent (who uses those artifacts to fulfill the request).
|
|
606
|
+
|
|
607
|
+
### Full structure
|
|
608
|
+
|
|
609
|
+
```ruby
|
|
610
|
+
module InterpretData
|
|
611
|
+
extend Workflow
|
|
612
|
+
self.include_workflow AgentWorkflow
|
|
613
|
+
|
|
614
|
+
# Phase 1: Analyst gathers and processes data into artifacts
|
|
615
|
+
chat_task :gather do
|
|
616
|
+
analyst = self.agent :Analyst, chat: chat, tooling: self.tooling
|
|
617
|
+
artifact_dir = file('artifacts')
|
|
618
|
+
|
|
619
|
+
# Dynamically inject a write_artifact task into the Analyst's workflow
|
|
620
|
+
analyst.workflow do
|
|
621
|
+
desc "Write an artifact to file"
|
|
622
|
+
input :name, :string, 'Name of the artifact', nil, required: true
|
|
623
|
+
input :content, :text, 'Content of the artifact', nil, required: true
|
|
624
|
+
task :write_artifact => :string do |name, content|
|
|
625
|
+
artifact_dir[name].write content
|
|
626
|
+
end
|
|
627
|
+
export_exec :write_artifact
|
|
628
|
+
end
|
|
629
|
+
|
|
630
|
+
analyst.user <<-EOF
|
|
631
|
+
Analyze the data and save data artifacts.
|
|
632
|
+
For text artifacts prefer the tool write_artifact.
|
|
633
|
+
You may also use other tools, like scripts, to create them but make sure
|
|
634
|
+
to save them in #{artifact_dir}.
|
|
635
|
+
Artifacts may include scripts for other agents to use to access the data.
|
|
636
|
+
|
|
637
|
+
Respond with a usage guide to these artifacts. Don't try to fulfill
|
|
638
|
+
completely the user request unless the final answer is obvious.
|
|
639
|
+
EOF
|
|
640
|
+
analyst
|
|
641
|
+
end
|
|
642
|
+
|
|
643
|
+
# Phase 2: Worker fulfills the request using gathered artifacts
|
|
644
|
+
dep :gather
|
|
645
|
+
chat_task :ask do
|
|
646
|
+
worker_agent = options[:InterpretData_worker_agent] || options[:worker_agent] || 'Worker'
|
|
647
|
+
chat = self.chat
|
|
648
|
+
chat.follow step(:gather).load
|
|
649
|
+
agent = self.agent worker_agent, chat: chat, tooling: self.tooling
|
|
650
|
+
agent.message :clear_tools, true
|
|
651
|
+
|
|
652
|
+
agent.user <<-EOF
|
|
653
|
+
Fulfill the request by using the artifacts:
|
|
654
|
+
|
|
655
|
+
#{step(:gather).file('artifacts').glob('**/**') * "\n" }
|
|
656
|
+
EOF
|
|
657
|
+
agent
|
|
658
|
+
end
|
|
659
|
+
end
|
|
660
|
+
```
|
|
661
|
+
|
|
662
|
+
### Key abstractions
|
|
663
|
+
|
|
664
|
+
1. **Dynamic workflow injection** — `analyst.workflow do ... end` defines a new task (`write_artifact`) at runtime and exposes it to the Analyst agent via `export_exec`. This lets the Analyst save files to a controlled directory.
|
|
665
|
+
|
|
666
|
+
2. **Artifact directory** — `file('artifacts')` creates a directory in the job's files area. Artifacts are persisted there and discoverable by subsequent stages.
|
|
667
|
+
|
|
668
|
+
3. **Artifact discovery** — the Worker stage lists all files in the artifact directory:
|
|
669
|
+
```ruby
|
|
670
|
+
step(:gather).file('artifacts').glob('**/**') * "\n"
|
|
671
|
+
```
|
|
672
|
+
This gives the Worker a file listing to work with, rather than embedding large data in the chat.
|
|
673
|
+
|
|
674
|
+
---
|
|
675
|
+
|
|
676
|
+
## The ChatAnalyst Pattern
|
|
677
|
+
|
|
678
|
+
### Overview
|
|
679
|
+
|
|
680
|
+
ChatAnalyst is a **meta-agent** — it inspects other agents' sessions rather than performing domain tasks. It is the most structurally complex workflow, with a `Session` class that traverses chat lineages, job dependencies, and agent logs.
|
|
681
|
+
|
|
682
|
+
### Key features
|
|
683
|
+
|
|
684
|
+
1. **Session discovery** — recursively follows chat references (`import`, `continue`, `last`) and job references to build a full graph of all related chats and jobs.
|
|
685
|
+
|
|
686
|
+
2. **Token accounting** — distinguishes between direct inference metadata (`pt`, `ct`, `tt`) and projection markers (`meta job=...`). Only direct metadata is counted; projections require following the job reference.
|
|
687
|
+
|
|
688
|
+
3. **Tool call analysis** — pairs `function_call`/`mcp_call` messages with `function_call_output` messages by call ID to determine success/failure.
|
|
689
|
+
|
|
690
|
+
4. **Agent interaction tracking** — identifies `ask` and `hand_off_to_*` calls specifically, useful for understanding delegation patterns.
|
|
691
|
+
|
|
692
|
+
5. **Exported tasks** — all tasks are `export_exec`, making them callable from CLI:
|
|
693
|
+
```ruby
|
|
694
|
+
export_exec :message_index, :message_content, :chat_overview,
|
|
695
|
+
:chat_tool_calls, :chat_tokens, :chat_agents, :chat_report
|
|
696
|
+
```
|
|
697
|
+
|
|
698
|
+
### Socialized chat files
|
|
699
|
+
|
|
700
|
+
ChatAnalyst documents the concept of **socialized chats** — projections of agent interactions:
|
|
701
|
+
|
|
702
|
+
> When a Manager or supervisor agent dispatches work to a specialist agent through the `ask` tool with a named `conversation`, the specialist interaction is persisted as a socialized chat file at:
|
|
703
|
+
>
|
|
704
|
+
> `<caller_job>.files/log/chats/<AgentName>/<conversation_name>.chat`
|
|
705
|
+
|
|
706
|
+
These are projections, not full logs. They contain the prompt, propagated options, a `meta: job=<path>` marker, and the assistant response — but carry zero direct inference tokens. The actual model calls are found by following the job reference.
|
|
707
|
+
|
|
708
|
+
---
|
|
709
|
+
|
|
710
|
+
## Delegation Flows
|
|
711
|
+
|
|
712
|
+
### How agents delegate to each other
|
|
713
|
+
|
|
714
|
+
There are two primary delegation mechanisms in SC26:
|
|
715
|
+
|
|
716
|
+
#### 1. Programmatic delegation (workflow-level)
|
|
717
|
+
|
|
718
|
+
Used by `Planned`, `Branched`, `Refined`, and `InterpretData`. The workflow code creates agents and drives them:
|
|
719
|
+
|
|
720
|
+
```ruby
|
|
721
|
+
agent = self.agent :Worker, chat: chat, tooling: self.tooling
|
|
722
|
+
agent.user "..."
|
|
723
|
+
agent.chat # blocks until the agent responds
|
|
724
|
+
```
|
|
725
|
+
|
|
726
|
+
This is synchronous, deterministic, and part of the Scout dependency graph. Each `chat_task` produces a cached job.
|
|
727
|
+
|
|
728
|
+
#### 2. Conversational delegation (Manager-level)
|
|
729
|
+
|
|
730
|
+
Used by the Manager. The Manager uses the `ask` tool during its live conversation:
|
|
731
|
+
|
|
732
|
+
```
|
|
733
|
+
socialize: true
|
|
734
|
+
tool: Manager
|
|
735
|
+
```
|
|
736
|
+
|
|
737
|
+
The Manager's `socialize: true` setting gives it access to all specialist agents. It calls `ask` with a target agent name, a prompt, and optionally a named `chat` identifier. This is asynchronous from the workflow perspective — the Manager decides at runtime whom to ask and what to say.
|
|
738
|
+
|
|
739
|
+
### `socialize` vs `delegate` usage
|
|
740
|
+
|
|
741
|
+
- **`socialize: true`** (in `start_chat`) — makes the agent able to see and interact with other agents. The Manager has this. It means the agent can use `ask` and `hand_off_to_*` tools.
|
|
742
|
+
- **No `socialize` / no `ask` tool** — agents like Worker and Critic cannot initiate delegation (unless explicitly given `ask` tools). The Critic has `no_ask_override: true` to explicitly prevent this.
|
|
743
|
+
|
|
744
|
+
Context: the Critic CAN ask the Worker questions (`"You have the ability to ask the Worker agent questions to clarify how or why things were done"`), but only for information, not for work.
|
|
745
|
+
|
|
746
|
+
### Context propagation patterns
|
|
747
|
+
|
|
748
|
+
| Pattern | Code | Effect |
|
|
749
|
+
|---------|------|--------|
|
|
750
|
+
| Full chat follow | `chat.follow step(:x).load` | All messages from stage x appended |
|
|
751
|
+
| Last message follow | `chat.follow step(:x).load.last` | Only the final answer from stage x |
|
|
752
|
+
| Dup for isolation | `chat: plan.dup` | Copy of plan chat, no shared mutation |
|
|
753
|
+
| Clear tools | `chat.message :clear_tools, true` | Strip tool definitions from context |
|
|
754
|
+
| Inline reference | `#{step(:x).load.answer}` | Embed answer text directly in a prompt |
|
|
755
|
+
|
|
756
|
+
### Named chat conversations
|
|
757
|
+
|
|
758
|
+
Named chats allow the Manager to maintain separate conversation threads:
|
|
759
|
+
|
|
760
|
+
- `plan_A`, `work_A`, `critic_A` — branch A's conversations
|
|
761
|
+
- `plan_B`, `work_B`, `critic_B` — branch B's conversations
|
|
762
|
+
|
|
763
|
+
Each named chat is a separate persisted conversation. The Manager can switch between them by specifying the `chat:` parameter in its `ask` calls. This is how budgeted branching works in practice: the Manager can pursue branch A, and if it fails, switch to branch B without losing either context.
|
|
764
|
+
|
|
765
|
+
---
|
|
766
|
+
|
|
767
|
+
## Reusable Patterns Summary
|
|
768
|
+
|
|
769
|
+
| Pattern | Structure | When to Use | Key Abstractions |
|
|
770
|
+
|---------|-----------|-------------|------------------|
|
|
771
|
+
| **Planned Pipeline** | `request → search → plan → work → ask` (linear dependency chain) | Well-defined tasks with clear phases; deterministic workflows | `dep`, `chat_task`, `chat.follow`, `task_alias`, conditional deps, configurable `worker_agent` |
|
|
772
|
+
| **Manager Control Loop** | `normalize → plan → score → execute → verify → (repair/switch/stop)` (model-driven loop) | Complex, uncertain tasks requiring adaptive decision-making; tasks with multiple solution strategies | `ask` tool, named chats, budget policy, structured delegation prompts, `socialize: true` |
|
|
773
|
+
| **Critic Verification** | `work → critic.evaluate → PASS/NEEDS_WORK/BLOCKED` | Any point where verification before continuation is needed | JSON schema response, `Chat.parse_json`, `no_ask_override`, evidence-based review, smallest-next-repair |
|
|
774
|
+
| **Branched Fan-Out** | `plan → spliter → [worker_1 ‖ worker_2 ‖ ...] → critic → report` | Embarrassingly parallel sub-tasks; same plan, different data | `iterate_dictionary`, `cpus: N`, `plan.dup` for isolation, unnamed spliter agent, aggregate Critic |
|
|
775
|
+
| **Refined Iteration** | `worker → critic → (NEEDS_WORK? retry) → PASS` | Tasks requiring quality convergence; iterative improvement | `TryAgain` exception, `retry`, shared chat across rounds, `clear_tools` between rounds |
|
|
776
|
+
| **InterpretData Prep** | `Analyst.gather(artifacts) → Worker.fulfill(artifacts)` | Tasks requiring data reduction before processing; large-volume data handling | Dynamic `analyst.workflow` injection, `export_exec`, artifact directory, `glob` for discovery |
|
|
777
|
+
| **ChatAnalyst Meta** | Session graph traversal → structured reports | Debugging agent sessions, analyzing token usage, understanding delegation patterns | `Session` class, recursive discovery, lineage tracking, `Chat.trace_chats`, `Chat.load` |
|
|
778
|
+
|
|
779
|
+
---
|
|
780
|
+
|
|
781
|
+
## Cross-Cutting Implementation Details
|
|
782
|
+
|
|
783
|
+
### `AgentWorkflow` inclusion
|
|
784
|
+
|
|
785
|
+
Every agent workflow module includes `AgentWorkflow`:
|
|
786
|
+
|
|
787
|
+
```ruby
|
|
788
|
+
module Planned
|
|
789
|
+
extend Workflow
|
|
790
|
+
self.include_workflow AgentWorkflow
|
|
791
|
+
```
|
|
792
|
+
|
|
793
|
+
This provides:
|
|
794
|
+
- `self.agent(name, chat:, tooling:)` — creates an agent instance with a given chat context and tooling.
|
|
795
|
+
- `self.chat` — access to the current chat.
|
|
796
|
+
- `self.tooling` / `self.tooling_intro` — tooling configuration for agents.
|
|
797
|
+
- `chat_task` — declares a task whose result is a chat (persisted as a `.chat` file).
|
|
798
|
+
|
|
799
|
+
### The `agent` factory method
|
|
800
|
+
|
|
801
|
+
```ruby
|
|
802
|
+
agent = self.agent :Worker, chat: chat, tooling: self.tooling
|
|
803
|
+
```
|
|
804
|
+
|
|
805
|
+
Parameters:
|
|
806
|
+
- **Agent name** (`:Worker`, `:Critic`, `:User`, `nil`) — selects the agent type. `nil` creates an unnamed agent.
|
|
807
|
+
- **`chat:`** — the chat context. Can be the current chat, a dup of another chat, or a named chat.
|
|
808
|
+
- **`tooling:`** — what tools to expose. `self.tooling` gives full tools; `self.tooling_intro` gives introductory/descriptive tooling; `false` gives no tools.
|
|
809
|
+
- **`no_ask_override:`** — when `true`, prevents the agent from getting `ask` delegation capabilities.
|
|
810
|
+
|
|
811
|
+
### `log_agent` for provenance
|
|
812
|
+
|
|
813
|
+
```ruby
|
|
814
|
+
log_agent worker, "worker-#{name}"
|
|
815
|
+
log_agent critic, "critic-round-#{round}"
|
|
816
|
+
```
|
|
817
|
+
|
|
818
|
+
This persists the agent's full chat log to the job's file area, enabling post-hoc inspection and the ChatAnalyst's analysis.
|
|
819
|
+
|
|
820
|
+
### Progress bars
|
|
821
|
+
|
|
822
|
+
```ruby
|
|
823
|
+
spliter.iterate_dictionary nil, cpus: 8, bar: self.progress_bar('Branches'), into: reports
|
|
824
|
+
```
|
|
825
|
+
|
|
826
|
+
`self.progress_bar('label')` creates a named progress bar for long-running parallel operations, providing visibility into execution progress.
|
|
827
|
+
|
|
828
|
+
### `IndiferentHash` for option access
|
|
829
|
+
|
|
830
|
+
```ruby
|
|
831
|
+
evaluation = IndiferentHash.setup(json)
|
|
832
|
+
evaluation[:status] # works with both string and symbol keys
|
|
833
|
+
```
|
|
834
|
+
|
|
835
|
+
This Scout utility makes hash access indifferent to whether keys are strings or symbols — essential when parsing JSON responses from agents.
|
|
836
|
+
|
|
837
|
+
---
|
|
838
|
+
|
|
839
|
+
## Pattern Selection Guide
|
|
840
|
+
|
|
841
|
+
| Situation | Recommended Pattern |
|
|
842
|
+
|-----------|-------------------|
|
|
843
|
+
| Task has clear phases (understand, research, plan, execute, report) | **Planned Pipeline** |
|
|
844
|
+
| Task is complex, uncertain, may need replanning or branch switching | **Manager Control Loop** |
|
|
845
|
+
| Task can be split into independent sub-tasks over different data | **Branched** |
|
|
846
|
+
| Task requires iterative quality improvement | **Refined** |
|
|
847
|
+
| Task involves large datasets that need reduction first | **InterpretData** |
|
|
848
|
+
| Need to verify work at any checkpoint | **Critic** (embed in any pattern) |
|
|
849
|
+
| Need to analyze past agent sessions | **ChatAnalyst** |
|
|
850
|
+
| Need to delegate to a specialist at runtime | **Manager with `ask`** |
|
|
851
|
+
| Need to delegate deterministically in a workflow | **Programmatic `self.agent`** |
|
|
852
|
+
|
|
853
|
+
These patterns can be composed: e.g., a Manager could delegate to a Branched workflow, which internally uses Refined for each branch, with a Critic at each level.
|