scout-ai 1.2.3 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.vimproject +138 -50
- data/README.md +171 -290
- data/Rakefile +17 -1
- data/VERSION +1 -1
- data/doc/Improvements.md +325 -0
- data/doc/StartHere.md +110 -0
- data/doc/developer/Architecture.md +126 -0
- data/doc/developer/Backends.md +199 -0
- data/doc/developer/ChatLifecycle.md +183 -0
- data/doc/developer/DelegationInternals.md +295 -0
- data/doc/developer/DesignPrinciples.md +245 -0
- data/doc/developer/PromptProcessing.md +292 -0
- data/doc/developer/Provenance.md +317 -0
- data/doc/user/BuildingAgents.md +345 -0
- data/doc/user/Cookbook.md +333 -0
- data/doc/user/CoreConcepts.md +181 -0
- data/doc/user/Delegation.md +191 -0
- data/doc/user/GettingStarted.md +159 -0
- data/doc/user/ManagingContext.md +163 -0
- data/doc/user/MultiAgentWorkflows.md +256 -0
- data/doc/user/Python.md +159 -0
- data/doc/user/RunningInference.md +200 -0
- data/doc/user/ToolCalling.md +193 -0
- data/doc/user/WritingChats.md +197 -0
- data/lib/scout/llm/agent/chat.rb +61 -11
- data/lib/scout/llm/agent/delegate.rb +274 -65
- data/lib/scout/llm/agent/iterate.rb +2 -2
- data/lib/scout/llm/agent/save.rb +273 -0
- data/lib/scout/llm/agent/workflow.rb +164 -0
- data/lib/scout/llm/agent.rb +86 -61
- data/lib/scout/llm/ask.rb +62 -17
- data/lib/scout/llm/backends/anthropic.rb +9 -2
- data/lib/scout/llm/backends/bedrock.rb +15 -3
- data/lib/scout/llm/backends/default.rb +183 -99
- data/lib/scout/llm/backends/glm.rb +58 -0
- data/lib/scout/llm/backends/huggingface.rb +196 -26
- data/lib/scout/llm/backends/ollama.rb +13 -1
- data/lib/scout/llm/backends/openai.rb +0 -2
- data/lib/scout/llm/backends/openwebui.rb +20 -13
- data/lib/scout/llm/backends/relay.rb +22 -22
- data/lib/scout/llm/backends/responses.rb +1 -1
- data/lib/scout/llm/chat/agent_meta.rb +264 -0
- data/lib/scout/llm/chat/annotation.rb +39 -10
- data/lib/scout/llm/chat/parse.rb +28 -6
- data/lib/scout/llm/chat/persist.rb +25 -0
- data/lib/scout/llm/chat/process/clear.rb +41 -6
- data/lib/scout/llm/chat/process/files.rb +21 -6
- data/lib/scout/llm/chat/process/meta.rb +421 -34
- data/lib/scout/llm/chat/process/options.rb +21 -1
- data/lib/scout/llm/chat/process/tools.rb +56 -15
- data/lib/scout/llm/chat/process.rb +4 -0
- data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
- data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
- data/lib/scout/llm/chat/prompt.rb +48 -0
- data/lib/scout/llm/chat/provenance.rb +775 -0
- data/lib/scout/llm/chat/tool_calls.rb +76 -0
- data/lib/scout/llm/chat.rb +18 -2
- data/lib/scout/llm/embed.rb +11 -3
- data/lib/scout/llm/image.rb +86 -0
- data/lib/scout/llm/mcp.rb +10 -2
- data/lib/scout/llm/rag.rb +3 -3
- data/lib/scout/llm/tools/call.rb +160 -11
- data/lib/scout/llm/tools/knowledge_base.rb +1 -1
- data/lib/scout/llm/tools/workflow.rb +32 -16
- data/lib/scout/model/python/huggingface/causal.rb +23 -5
- data/lib/scout/model/python/huggingface.rb +2 -1
- data/lib/scout-ai.rb +1 -0
- data/python/README.md +197 -14
- data/python/scout_ai/huggingface/eval.py +245 -34
- data/python/tests/test_huggingface_eval.py +58 -0
- data/research/ChatAnalyst-required-changes.md +167 -0
- data/research/agent-delegation-analysis.md +810 -0
- data/research/agent-meta-provenance-integration-plan.md +622 -0
- data/research/agent-workflow-analysis.md +1120 -0
- data/research/backends-analysis.md +836 -0
- data/research/chat-core-analysis.md +946 -0
- data/research/chatanalyst-provenance/00-baseline.md +30 -0
- data/research/chatanalyst-provenance/01-repo-map.md +60 -0
- data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
- data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
- data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
- data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
- data/research/chatanalyst-provenance/07-critic-review.md +25 -0
- data/research/chatanalyst-provenance/final-report.md +45 -0
- data/research/chatanalyst-provenance/resumption.md +37 -0
- data/research/coding-philosophy-analysis.md +928 -0
- data/research/commands-analysis.md +947 -0
- data/research/multi-agent-patterns-analysis.md +853 -0
- data/research/prompt-strategies-analysis.md +630 -0
- data/research/prov-verbosity-fix-notes.md +77 -0
- data/research/provenance-analysis.md +469 -0
- data/research/provenance-navigation-design.md +640 -0
- data/research/synthesis-report.md +487 -0
- data/research/tools-system-analysis.md +779 -0
- data/scout-ai.gemspec +100 -11
- data/scout_commands/agent/ask +13 -3
- data/scout_commands/agent/kb +2 -0
- data/scout_commands/llm/ask +11 -4
- data/scout_commands/llm/md +76 -0
- data/scout_commands/llm/process_queries +48 -0
- data/scout_commands/llm/prov +602 -0
- data/scout_commands/llm/word +71 -0
- data/scout_commands/workflow/mcp +43 -0
- data/share/word/reference.docx +0 -0
- data/test/etc/AI/mock.yaml +11 -0
- data/test/fixtures/backends/anthropic.json +19 -0
- data/test/fixtures/backends/anthropic_tool_use.json +24 -0
- data/test/fixtures/backends/bedrock.json +8 -0
- data/test/fixtures/backends/bedrock_embedding.json +3 -0
- data/test/fixtures/backends/bedrock_tool_use.json +17 -0
- data/test/fixtures/backends/ollama.json +16 -0
- data/test/fixtures/backends/ollama_tool_call.json +27 -0
- data/test/fixtures/backends/openai_chat.json +21 -0
- data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
- data/test/fixtures/backends/responses.json +33 -0
- data/test/fixtures/backends/responses_tool_call.json +28 -0
- data/test/integration/README.md +32 -0
- data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
- data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
- data/test/integration/scout/llm/backends/test_relay.rb +52 -0
- data/test/integration/scout/llm/test_infrastructure.rb +74 -0
- data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
- data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
- data/test/integration/scout/model/test_base.rb +91 -0
- data/test/scout/llm/agent/test_chat.rb +8 -2
- data/test/scout/llm/agent/test_save.rb +413 -0
- data/test/scout/llm/agent/test_workflow.rb +110 -0
- data/test/scout/llm/backends/test_anthropic.rb +93 -10
- data/test/scout/llm/backends/test_bedrock.rb +118 -2
- data/test/scout/llm/backends/test_huggingface.rb +137 -42
- data/test/scout/llm/backends/test_ollama.rb +70 -20
- data/test/scout/llm/backends/test_openwebui.rb +42 -40
- data/test/scout/llm/backends/test_relay.rb +4 -2
- data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
- data/test/scout/llm/chat/process/test_meta.rb +518 -0
- data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
- data/test/scout/llm/chat/test_agent_meta.rb +357 -0
- data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
- data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
- data/test/scout/llm/chat/test_parse.rb +70 -15
- data/test/scout/llm/chat/test_prov_cli.rb +274 -0
- data/test/scout/llm/chat/test_provenance.rb +240 -0
- data/test/scout/llm/chat/test_tool_calls.rb +38 -0
- data/test/scout/llm/test_agent.rb +13 -36
- data/test/scout/llm/test_ask.rb +75 -52
- data/test/scout/llm/test_chat.rb +107 -13
- data/test/scout/llm/test_embed.rb +48 -0
- data/test/scout/llm/test_rag.rb +23 -16
- data/test/scout/llm/test_tools.rb +12 -1
- data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
- data/test/scout/llm/tools/test_mcp.rb +5 -3
- data/test/scout/llm/tools/test_workflow.rb +23 -2
- data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
- data/test/scout/model/python/huggingface/test_causal.rb +9 -3
- data/test/scout/model/python/huggingface/test_classification.rb +11 -2
- data/test/scout/model/python/test_torch.rb +2 -0
- data/test/scout/model/python/torch/test_helpers.rb +4 -0
- data/test/scout/model/test_base.rb +4 -2
- data/test/support/availability.rb +231 -0
- data/test/support/fake_clients.rb +138 -0
- data/test/support/fixtures.rb +21 -0
- data/test/support/infrastructure_probes.rb +136 -0
- data/test/support/mock_backend.rb +215 -0
- data/test/test_helper.rb +32 -2
- metadata +99 -10
- data/doc/Agent.md +0 -327
- data/doc/Chat.md +0 -458
- data/doc/LLM.md +0 -340
- data/doc/RAG.md +0 -129
- data/scout_commands/documenter +0 -148
- data/test/scout/llm/backends/test_openai.rb +0 -192
- data/test/scout/llm/backends/test_responses.rb +0 -238
- data/test/scout/llm/test_parse.rb +0 -98
|
@@ -0,0 +1,245 @@
|
|
|
1
|
+
# Design Principles
|
|
2
|
+
|
|
3
|
+
This document explains the coding philosophy and idioms that make Scout-AI
|
|
4
|
+
elegant and expressive. It is intended for framework contributors who want to
|
|
5
|
+
write code that fits the existing style.
|
|
6
|
+
|
|
7
|
+
> For detailed examples and analysis of idiomatic vs. non-idiomatic patterns,
|
|
8
|
+
> see [../../research/coding-philosophy-analysis.md](../../research/coding-philosophy-analysis.md).
|
|
9
|
+
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
## Abstraction-first
|
|
13
|
+
|
|
14
|
+
Every concept in Scout-AI is an abstraction with a crisp boundary:
|
|
15
|
+
|
|
16
|
+
- A **Chat** is "a conversation."
|
|
17
|
+
- An **Agent** is "a stateful conversation holder with tools."
|
|
18
|
+
- A **Backend** is "an adapter to a model API."
|
|
19
|
+
- A **Tool** is "a callable function exposed to the model."
|
|
20
|
+
|
|
21
|
+
New features should be expressed as a new abstraction or an extension of an
|
|
22
|
+
existing one, not as inline logic scattered across files. The question is
|
|
23
|
+
always: *what abstraction does this feature belong to?*
|
|
24
|
+
|
|
25
|
+
---
|
|
26
|
+
|
|
27
|
+
## Module composition over inheritance
|
|
28
|
+
|
|
29
|
+
Scout-AI avoids deep class hierarchies. Behavior is composed through Ruby
|
|
30
|
+
modules:
|
|
31
|
+
|
|
32
|
+
```ruby
|
|
33
|
+
# Agent's behavior is split across multiple files that reopen the class:
|
|
34
|
+
# lib/scout/llm/agent.rb — core (ask, prompt, workflow)
|
|
35
|
+
# lib/scout/llm/agent/chat.rb — chat management
|
|
36
|
+
# lib/scout/llm/agent/iterate.rb — iteration patterns
|
|
37
|
+
# lib/scout/llm/agent/delegate.rb — multi-agent delegation
|
|
38
|
+
# lib/scout/llm/agent/workflow.rb — AgentWorkflow mixin + chat_task DSL
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
Each module adds a cohesive set of methods to the same class. There is no
|
|
42
|
+
inheritance tree — just flat composition. This keeps each concern in its own
|
|
43
|
+
file while sharing `@other_options`, `@current_chat`, etc.
|
|
44
|
+
|
|
45
|
+
---
|
|
46
|
+
|
|
47
|
+
## Chat-as-data (annotate, don't wrap)
|
|
48
|
+
|
|
49
|
+
The single most important design decision:
|
|
50
|
+
|
|
51
|
+
> **A Chat is a plain `Array` of message `Hash`es, not an opaque object.**
|
|
52
|
+
|
|
53
|
+
The `Chat` module uses scout-essentials' `Annotation` system to add DSL methods
|
|
54
|
+
to a plain Array **non-invasively**:
|
|
55
|
+
|
|
56
|
+
```ruby
|
|
57
|
+
chat = Chat.setup([])
|
|
58
|
+
chat.user("Hello")
|
|
59
|
+
|
|
60
|
+
chat.class # => Array (still an Array!)
|
|
61
|
+
chat.first[:role] # => "user"
|
|
62
|
+
chat.select { |m| m[:role] == 'system' } # standard Array operations work
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
The annotation:
|
|
66
|
+
- Adds methods to the **singleton class** of the specific object instance.
|
|
67
|
+
- Does **not** change the object's class.
|
|
68
|
+
- Is **removable** via `Annotation.purge(obj)`.
|
|
69
|
+
|
|
70
|
+
**Implication:** Don't create wrapper classes for data that is already a Hash
|
|
71
|
+
or Array. Annotate it instead.
|
|
72
|
+
|
|
73
|
+
---
|
|
74
|
+
|
|
75
|
+
## Convention over configuration
|
|
76
|
+
|
|
77
|
+
Scout-AI discovers components by convention rather than registration:
|
|
78
|
+
|
|
79
|
+
| Convention | Resolution |
|
|
80
|
+
|---|---|
|
|
81
|
+
| Agent directory `Agent/<Name>/` | Auto-discovered via `Scout.Agent`, `Scout.chats.Agent`, etc. |
|
|
82
|
+
| `start_chat` file | Loaded as initial conversation. |
|
|
83
|
+
| `workflow.rb` | Loaded as the agent's workflow. |
|
|
84
|
+
| `knowledge_base/` | Loaded as the agent's KB. |
|
|
85
|
+
| `python/*.py` | Loaded as Python-backed tools. |
|
|
86
|
+
|
|
87
|
+
There are no registration calls or plugin manifests. Put files in the right
|
|
88
|
+
place and they are found.
|
|
89
|
+
|
|
90
|
+
---
|
|
91
|
+
|
|
92
|
+
## The DSL pattern
|
|
93
|
+
|
|
94
|
+
Scout-AI builds expressive domain-specific languages on top of Ruby's
|
|
95
|
+
flexibility:
|
|
96
|
+
|
|
97
|
+
### Chat DSL
|
|
98
|
+
|
|
99
|
+
```ruby
|
|
100
|
+
chat = Chat.setup([])
|
|
101
|
+
chat.system("You are a helpful assistant")
|
|
102
|
+
chat.user("What is 2+2?")
|
|
103
|
+
chat.option(:model, "gpt-4")
|
|
104
|
+
chat.ask # → "4"
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
### Agent DSL (via method_missing proxy)
|
|
108
|
+
|
|
109
|
+
```ruby
|
|
110
|
+
agent = LLM.agent(model: "gpt-4")
|
|
111
|
+
agent.system("You are a coder")
|
|
112
|
+
agent.user("Write a function")
|
|
113
|
+
agent.chat
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
The `method_missing` proxy on Agent forwards unknown methods to
|
|
117
|
+
`current_chat`, so all Chat methods work directly on the Agent.
|
|
118
|
+
|
|
119
|
+
### Workflow DSL (chat_task)
|
|
120
|
+
|
|
121
|
+
```ruby
|
|
122
|
+
module MyPipeline
|
|
123
|
+
extend Workflow
|
|
124
|
+
self.include_workflow AgentWorkflow
|
|
125
|
+
|
|
126
|
+
chat_task :my_task do
|
|
127
|
+
agent = self.agent :Worker, chat: chat
|
|
128
|
+
agent.user "Do something"
|
|
129
|
+
agent
|
|
130
|
+
end
|
|
131
|
+
end
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
---
|
|
135
|
+
|
|
136
|
+
## Lazy initialization
|
|
137
|
+
|
|
138
|
+
Many things are initialized on first use, not eagerly:
|
|
139
|
+
|
|
140
|
+
```ruby
|
|
141
|
+
def workflow(&block)
|
|
142
|
+
@workflow ||= begin
|
|
143
|
+
m = Module.new
|
|
144
|
+
m.extend Workflow
|
|
145
|
+
m.name ||= 'ScoutAgent'
|
|
146
|
+
m.tasks = {}
|
|
147
|
+
m
|
|
148
|
+
end
|
|
149
|
+
end
|
|
150
|
+
|
|
151
|
+
def current_chat
|
|
152
|
+
@current_chat ||= start
|
|
153
|
+
end
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
This keeps object creation cheap and defers expensive setup until needed.
|
|
157
|
+
|
|
158
|
+
---
|
|
159
|
+
|
|
160
|
+
## IndiferentHash everywhere
|
|
161
|
+
|
|
162
|
+
Options and metadata use Scout's `IndiferentHash` (symbol/string-indifferent
|
|
163
|
+
access):
|
|
164
|
+
|
|
165
|
+
```ruby
|
|
166
|
+
options[:model] # works
|
|
167
|
+
options['model'] # also works
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
This eliminates a whole class of symbol-vs-string bugs. When building option
|
|
171
|
+
hashes, use `IndiferentHash.setup(hash)` rather than a plain Hash.
|
|
172
|
+
|
|
173
|
+
---
|
|
174
|
+
|
|
175
|
+
## Idiomatic patterns to follow
|
|
176
|
+
|
|
177
|
+
### Use `Chat.setup` not `Chat.new`
|
|
178
|
+
|
|
179
|
+
```ruby
|
|
180
|
+
# Good
|
|
181
|
+
chat = Chat.setup([])
|
|
182
|
+
|
|
183
|
+
# Wrong — Chat is a module, not a class
|
|
184
|
+
chat = Chat.new # NoMethodError
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
### Prefer annotation forwarding over explicit wrappers
|
|
188
|
+
|
|
189
|
+
```ruby
|
|
190
|
+
# Good — Agent forwards to Chat via method_missing
|
|
191
|
+
agent.user("hello")
|
|
192
|
+
|
|
193
|
+
# Wrong — don't write explicit delegation methods
|
|
194
|
+
def agent_user(agent, msg)
|
|
195
|
+
agent.current_chat.user(msg)
|
|
196
|
+
end
|
|
197
|
+
```
|
|
198
|
+
|
|
199
|
+
### Use `Chat.follow` to compose conversations
|
|
200
|
+
|
|
201
|
+
```ruby
|
|
202
|
+
# Good — follow prepends context
|
|
203
|
+
chat.follow(step(:plan).load)
|
|
204
|
+
|
|
205
|
+
# Avoid — manual concatenation loses annotations
|
|
206
|
+
chat = step(:plan).load + chat
|
|
207
|
+
```
|
|
208
|
+
|
|
209
|
+
### Prefer `proc` blocks for tools
|
|
210
|
+
|
|
211
|
+
```ruby
|
|
212
|
+
# Good — Proc-based tool
|
|
213
|
+
LLM.add_tool(
|
|
214
|
+
"my_tool",
|
|
215
|
+
"Does something useful",
|
|
216
|
+
{"type" => "object", "properties" => {...}}
|
|
217
|
+
) do |name, params|
|
|
218
|
+
# tool implementation
|
|
219
|
+
end
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
---
|
|
223
|
+
|
|
224
|
+
## Anti-patterns to avoid
|
|
225
|
+
|
|
226
|
+
1. **Creating wrapper classes for Arrays/Hashes** — Annotate instead.
|
|
227
|
+
2. **Adding provider-specific logic to `LLM.ask`** — Put it in the Backend module.
|
|
228
|
+
3. **Deep inheritance hierarchies** — Use module composition.
|
|
229
|
+
4. **Eager initialization** — Use lazy `||=`.
|
|
230
|
+
5. **Explicit delegation methods when `method_missing` already works** — Agent
|
|
231
|
+
already forwards to Chat; don't add `agent_user`, `agent_system`, etc.
|
|
232
|
+
6. **Mutating the stored chat during prompt preparation** — Prompt strategies
|
|
233
|
+
are ephemeral; never mutate the source.
|
|
234
|
+
7. **Using `Marshal.dump/load` for deep copying** — Procs can't be marshalled;
|
|
235
|
+
use `social_duplicate` patterns.
|
|
236
|
+
8. **Hard-coding agent names in delegation logic** — Let the model choose via
|
|
237
|
+
`socialize`, or use `delegate` with explicit instances.
|
|
238
|
+
|
|
239
|
+
---
|
|
240
|
+
|
|
241
|
+
## Cross-references
|
|
242
|
+
|
|
243
|
+
- [Architecture.md](Architecture.md) — Overall system architecture.
|
|
244
|
+
- [ChatLifecycle.md](ChatLifecycle.md) — Chat data model.
|
|
245
|
+
- [../../research/coding-philosophy-analysis.md](../../research/coding-philosophy-analysis.md) — Deep investigation.
|
|
@@ -0,0 +1,292 @@
|
|
|
1
|
+
# Prompt Processing
|
|
2
|
+
|
|
3
|
+
This document explains the internal mechanism Scout-AI uses to manage long
|
|
4
|
+
contexts before sending a prompt to the LLM. It is intended for framework
|
|
5
|
+
contributors.
|
|
6
|
+
|
|
7
|
+
> For the user-facing guide on what happens when contexts get long, see
|
|
8
|
+
> [../user/ManagingContext.md](../user/ManagingContext.md).
|
|
9
|
+
> For deep code investigation, see
|
|
10
|
+
> [../../research/prompt-strategies-analysis.md](../../research/prompt-strategies-analysis.md).
|
|
11
|
+
|
|
12
|
+
---
|
|
13
|
+
|
|
14
|
+
## What prompt strategies are
|
|
15
|
+
|
|
16
|
+
Prompt strategies are a **pre-inference transformation layer** that modifies
|
|
17
|
+
the chat message list *just before* it is sent to the LLM backend. The primary
|
|
18
|
+
motivation is context-window management: in long agent conversations with many
|
|
19
|
+
tool calls, the accumulated arguments and return values can consume enormous
|
|
20
|
+
amounts of tokens.
|
|
21
|
+
|
|
22
|
+
The system works by applying named "strategies" to the message array. Each
|
|
23
|
+
strategy is a function that takes an Array of message hashes and returns a
|
|
24
|
+
(possibly shorter or modified) Array.
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
## File layout
|
|
29
|
+
|
|
30
|
+
Strategy implementations live in `lib/scout/llm/prompt/`:
|
|
31
|
+
|
|
32
|
+
```
|
|
33
|
+
lib/scout/llm/
|
|
34
|
+
├── chat/
|
|
35
|
+
│ └── prompt.rb # Dispatcher: prepare_prompt, shared constants
|
|
36
|
+
└── prompt/
|
|
37
|
+
├── shorten_tools.rb # Default strategy (recomputes each turn)
|
|
38
|
+
└── shorten_tools_epoch.rb # Cache-friendly epoch variant
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
The dispatcher in `prompt.rb` requires both strategy files and delegates to
|
|
42
|
+
them via a `case` statement inside `prepare_prompt`.
|
|
43
|
+
|
|
44
|
+
---
|
|
45
|
+
|
|
46
|
+
## The ephemeral design
|
|
47
|
+
|
|
48
|
+
**Prompt strategies never mutate the stored chat.** The transformation happens
|
|
49
|
+
entirely inside the backend's `ask` method:
|
|
50
|
+
|
|
51
|
+
```ruby
|
|
52
|
+
# lib/scout/llm/backends/default.rb
|
|
53
|
+
prompt = Chat.prepare_prompt(messages, prompt_strategies)
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
The variable `prompt` is a local derived from `messages`. The original messages
|
|
57
|
+
and the underlying Chat Object are untouched. This means:
|
|
58
|
+
|
|
59
|
+
- The chat history retains full-fidelity tool outputs for later inspection.
|
|
60
|
+
- The agent's persisted memory is not degraded by truncation.
|
|
61
|
+
- Only the *next inference* sees the shortened prompt.
|
|
62
|
+
|
|
63
|
+
This deliberately decouples *what the model sees* from *what the system
|
|
64
|
+
remembers*.
|
|
65
|
+
|
|
66
|
+
---
|
|
67
|
+
|
|
68
|
+
## `prepare_prompt` entry point
|
|
69
|
+
|
|
70
|
+
```ruby
|
|
71
|
+
def self.prepare_prompt(prompt, prompt_strategies = nil)
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
The method supports four input forms for `prompt_strategies`:
|
|
75
|
+
|
|
76
|
+
| Input type | Behavior |
|
|
77
|
+
|---|---|
|
|
78
|
+
| `Proc` | Called directly with the prompt array — full custom hook. |
|
|
79
|
+
| `nil` | Falls back to `DEFAULT_CONTEXT_STRATEGY` = `%w(shorten_tools)`. |
|
|
80
|
+
| `String` | Split by comma into strategy names (e.g., `"shorten_tools,custom"`). |
|
|
81
|
+
| `Array<String>` | Apply each named strategy in sequence. |
|
|
82
|
+
|
|
83
|
+
Strategies are applied **in sequence**: each receives the output of the previous.
|
|
84
|
+
|
|
85
|
+
The string `"none"` is a recognized no-op that returns the prompt unchanged.
|
|
86
|
+
|
|
87
|
+
---
|
|
88
|
+
|
|
89
|
+
## The `shorten_tools` strategy (default)
|
|
90
|
+
|
|
91
|
+
This is the **default strategy** — it runs on every backend `ask` call unless
|
|
92
|
+
explicitly disabled.
|
|
93
|
+
|
|
94
|
+
### Algorithm
|
|
95
|
+
|
|
96
|
+
`shorten_tools` walks the message array **in reverse** (newest-first) and
|
|
97
|
+
applies a tiered truncation/dropping policy to `function_call` and
|
|
98
|
+
`function_call_output` messages. The most recent tool interactions get priority
|
|
99
|
+
for full retention; older ones are progressively degraded.
|
|
100
|
+
|
|
101
|
+
Three counters are tracked during the reverse traversal:
|
|
102
|
+
|
|
103
|
+
| Counter | Meaning |
|
|
104
|
+
|---|---|
|
|
105
|
+
| `tool_ids` | Count of tool output messages encountered (from the end) |
|
|
106
|
+
| `tool_chars` | Cumulative characters of retained tool content |
|
|
107
|
+
| `user_messages` | Number of user messages encountered |
|
|
108
|
+
|
|
109
|
+
### Three-tier degradation
|
|
110
|
+
|
|
111
|
+
For tool outputs (the same logic applies to tool calls with separate thresholds):
|
|
112
|
+
|
|
113
|
+
| Position (from end) | Condition | Action |
|
|
114
|
+
|---|---|---|
|
|
115
|
+
| Most recent N | `count < full_tool_outputs` | **Full fidelity** |
|
|
116
|
+
| Middle band | Between `full_tool_outputs` and `max_tool_outputs` | **Truncated** (content shortened, hash-stamped) |
|
|
117
|
+
| Beyond max | `count > max_tool_outputs` | **Dropped entirely** |
|
|
118
|
+
|
|
119
|
+
There is also a character-budget override: if cumulative tool content is below
|
|
120
|
+
`max_tool_chars`, messages are kept at full fidelity regardless of position.
|
|
121
|
+
This means **short conversations are never truncated** — the system is a no-op
|
|
122
|
+
until context pressure is real.
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## The `shorten_tools_epoch` strategy (cache-friendly)
|
|
127
|
+
|
|
128
|
+
### Motivation
|
|
129
|
+
|
|
130
|
+
The default `shorten_tools` strategy recomputes the truncation boundary on
|
|
131
|
+
**every single inference**. When a new tool call is added, the boundary shifts
|
|
132
|
+
by one position, causing every previously-truncated message to be re-evaluated
|
|
133
|
+
with a different offset. This means the prompt prefix changes on every turn,
|
|
134
|
+
**defeating KV-cache and prompt-cache mechanisms** offered by LLM providers.
|
|
135
|
+
|
|
136
|
+
`shorten_tools_epoch` solves this by **freezing the compaction boundary** for
|
|
137
|
+
windows of N tool calls called *epochs*. Within an epoch, the compacted prefix
|
|
138
|
+
is byte-for-byte identical across consecutive inferences, maximizing cache hit
|
|
139
|
+
rates.
|
|
140
|
+
|
|
141
|
+
### Algorithm
|
|
142
|
+
|
|
143
|
+
The conversation is divided into four regions (newest at the bottom):
|
|
144
|
+
|
|
145
|
+
```
|
|
146
|
+
[ dropped ] tool calls older than (compacted + full) → removed entirely
|
|
147
|
+
[ compacted ] up to epoch_compacted_tool_calls tool calls, truncated
|
|
148
|
+
[ full-recent ] epoch_full_tool_calls tool calls at full fidelity
|
|
149
|
+
[ full-new ] any tool calls that arrived after the epoch boundary (full fidelity)
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
The `compacted` and `full-recent` regions are **pinned** relative to the epoch
|
|
153
|
+
boundary, not the live tool-call count. Their content stays stable until the
|
|
154
|
+
boundary advances.
|
|
155
|
+
|
|
156
|
+
### Epoch boundary calculation
|
|
157
|
+
|
|
158
|
+
```
|
|
159
|
+
overflow = total_tool_calls - threshold # how many beyond threshold
|
|
160
|
+
epoch_idx = overflow > 0 ? (overflow - 1) / epoch_size : 0
|
|
161
|
+
pinned_total = threshold + (epoch_idx * epoch_size)
|
|
162
|
+
new_calls = total_tool_calls - pinned_total # tool calls that arrived this epoch
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
The `pinned_total` determines where the `full-recent` region starts. Any tool
|
|
166
|
+
calls beyond `pinned_total` are treated as "new" and kept at full fidelity.
|
|
167
|
+
|
|
168
|
+
### Worked example (threshold=100, full=10, compacted=40, epoch_size=10)
|
|
169
|
+
|
|
170
|
+
| Tool calls | pinned_total | new_calls | keep_full | compacted | dropped |
|
|
171
|
+
|---|---|---|---|---|---|
|
|
172
|
+
| 100 | 100 | 0 | 10 | 40 | 50 |
|
|
173
|
+
| 101 | 100 | 1 | 11 | 40 | 50 |
|
|
174
|
+
| 105 | 100 | 5 | 15 | 40 | 50 |
|
|
175
|
+
| 110 | 100 | 10 | 20 | 40 | 50 |
|
|
176
|
+
| 111 | 110 | 1 | 11 | 40 | 60 |
|
|
177
|
+
|
|
178
|
+
From tool calls 101–110 the compacted region (calls 11–50 from the pinned
|
|
179
|
+
boundary) is **identical**, so the prompt prefix is cache-stable for 10
|
|
180
|
+
consecutive inferences. At call 111 the boundary advances and the compacted
|
|
181
|
+
region shifts.
|
|
182
|
+
|
|
183
|
+
### Configuration
|
|
184
|
+
|
|
185
|
+
| Config key | ENV var | Default | Description |
|
|
186
|
+
|---|---|---|---|
|
|
187
|
+
| `epoch_tool_call_threshold` | `EPOCH_TOOL_CALL_THRESHOLD` | 50 | Total tool calls at or below which no compaction happens |
|
|
188
|
+
| `epoch_full_tool_calls` | `EPOCH_FULL_TOOL_CALLS` | 10 | Most-recent tool calls kept at full fidelity |
|
|
189
|
+
| `epoch_compacted_tool_calls` | `EPOCH_COMPACTED_TOOL_CALLS` | 40 | Tool calls (before full-recent) to truncate |
|
|
190
|
+
| `epoch_size` | `EPOCH_SIZE` | 10 | New tool calls allowed before boundary advances |
|
|
191
|
+
|
|
192
|
+
All thresholds are read via `Scout::Config.get` and memoized in class variables,
|
|
193
|
+
following the same pattern as `shorten_tools`.
|
|
194
|
+
|
|
195
|
+
### Enabling the epoch strategy
|
|
196
|
+
|
|
197
|
+
To use it instead of the default, pass the strategy name:
|
|
198
|
+
|
|
199
|
+
```ruby
|
|
200
|
+
# In options
|
|
201
|
+
options[:prompt_strategies] = 'shorten_tools_epoch'
|
|
202
|
+
|
|
203
|
+
# Or directly
|
|
204
|
+
Chat.prepare_prompt(messages, 'shorten_tools_epoch')
|
|
205
|
+
```
|
|
206
|
+
|
|
207
|
+
To switch the default system-wide, set:
|
|
208
|
+
|
|
209
|
+
```ruby
|
|
210
|
+
Chat::DEFAULT_CONTEXT_STRATEGY.replace(['shorten_tools_epoch'])
|
|
211
|
+
```
|
|
212
|
+
|
|
213
|
+
---
|
|
214
|
+
|
|
215
|
+
## Configuration thresholds (`shorten_tools`)
|
|
216
|
+
|
|
217
|
+
All thresholds are read via `Scout::Config.get` and **memoized** in class
|
|
218
|
+
variables on first access:
|
|
219
|
+
|
|
220
|
+
| Config key | ENV var | Default | Description |
|
|
221
|
+
|---|---|---|---|
|
|
222
|
+
| `full_tool_calls` | `FULL_TOOL_CALLS` | 0 | Recent tool calls kept at full fidelity |
|
|
223
|
+
| `full_tool_outputs` | `FULL_TOOL_OUTPUTS` | 10 | Recent tool outputs kept at full fidelity |
|
|
224
|
+
| `max_tool_calls` | `MAX_TOOL_CALLS` | 40 | Hard limit; tool calls beyond this are dropped |
|
|
225
|
+
| `max_tool_outputs` | `MAX_TOOL_OUTPUTS` | 40 (defaults to `max_tool_calls`) | Hard limit; outputs beyond this are dropped |
|
|
226
|
+
| `max_tool_chars` | `MAX_TOOL_CHARS` | 100,000 | Cumulative character budget for retained tool content |
|
|
227
|
+
|
|
228
|
+
### Memoization trade-off
|
|
229
|
+
|
|
230
|
+
Because thresholds use `||=` memoization, they are **frozen for the process
|
|
231
|
+
lifetime** after first access. Changing config files or ENV vars mid-process
|
|
232
|
+
has no effect. This is fine for CLI/agent usage but could be surprising in
|
|
233
|
+
long-running daemons.
|
|
234
|
+
|
|
235
|
+
---
|
|
236
|
+
|
|
237
|
+
## Content hashing in truncated strings
|
|
238
|
+
|
|
239
|
+
When a value is truncated, `Log.truncate_string` embeds an MD5 hash prefix:
|
|
240
|
+
|
|
241
|
+
```
|
|
242
|
+
Truncated (15432): The first ~70 chars...<...15432 - a1b2c...>...last ~70 chars
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
This allows truncated content to be matched against logs or the original chat
|
|
246
|
+
for debugging.
|
|
247
|
+
|
|
248
|
+
---
|
|
249
|
+
|
|
250
|
+
## Integration with the backend
|
|
251
|
+
|
|
252
|
+
`prepare_prompt` is called inside `Backend::Default#ask`, in the normal
|
|
253
|
+
(non-relay) path:
|
|
254
|
+
|
|
255
|
+
```ruby
|
|
256
|
+
client = prepare_client(options, messages)
|
|
257
|
+
prompt = Chat.prepare_prompt(messages, prompt_strategies)
|
|
258
|
+
formatted_prompt = format_messages(prompt)
|
|
259
|
+
tools = tools(formatted_prompt, options)
|
|
260
|
+
response = query(client, formatted_prompt, tools, options)
|
|
261
|
+
```
|
|
262
|
+
|
|
263
|
+
Key points:
|
|
264
|
+
- Applied on every `ask` invocation in the normal path.
|
|
265
|
+
- **Not applied in relay mode** (raw messages are uploaded to a remote server).
|
|
266
|
+
- Applied **before** `format_messages` — strategy output directly determines
|
|
267
|
+
token consumption.
|
|
268
|
+
- `prompt_strategies` comes from the `options` hash, so callers can override
|
|
269
|
+
per-call.
|
|
270
|
+
|
|
271
|
+
---
|
|
272
|
+
|
|
273
|
+
## Extension point: custom strategies
|
|
274
|
+
|
|
275
|
+
Two mechanisms coexist:
|
|
276
|
+
|
|
277
|
+
1. **Hard-coded `case` dispatch** for built-in strategies (`shorten_tools`,
|
|
278
|
+
`shorten_tools_epoch`, `none`).
|
|
279
|
+
2. **`REGISTERED_STRATEGIES` hash** for user/plugin-registered strategies.
|
|
280
|
+
|
|
281
|
+
> **Note:** `REGISTERED_STRATEGIES` is referenced in the code but not yet
|
|
282
|
+
> populated with entries. Passing an unknown strategy name will result in
|
|
283
|
+
> `nil.call(prompt)`, raising a `NoMethodError`. For now, use a `Proc` to
|
|
284
|
+
> supply custom strategies.
|
|
285
|
+
|
|
286
|
+
---
|
|
287
|
+
|
|
288
|
+
## Cross-references
|
|
289
|
+
|
|
290
|
+
- [../user/ManagingContext.md](../user/ManagingContext.md) — User guide for long contexts.
|
|
291
|
+
- [Backends.md](Backends.md) — Where `prepare_prompt` is called in the inference loop.
|
|
292
|
+
- [../../research/prompt-strategies-analysis.md](../../research/prompt-strategies-analysis.md) — Deep investigation.
|