scout-ai 1.2.3 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (174) hide show
  1. checksums.yaml +4 -4
  2. data/.vimproject +138 -50
  3. data/README.md +171 -290
  4. data/Rakefile +17 -1
  5. data/VERSION +1 -1
  6. data/doc/Improvements.md +325 -0
  7. data/doc/StartHere.md +110 -0
  8. data/doc/developer/Architecture.md +126 -0
  9. data/doc/developer/Backends.md +199 -0
  10. data/doc/developer/ChatLifecycle.md +183 -0
  11. data/doc/developer/DelegationInternals.md +295 -0
  12. data/doc/developer/DesignPrinciples.md +245 -0
  13. data/doc/developer/PromptProcessing.md +292 -0
  14. data/doc/developer/Provenance.md +317 -0
  15. data/doc/user/BuildingAgents.md +345 -0
  16. data/doc/user/Cookbook.md +333 -0
  17. data/doc/user/CoreConcepts.md +181 -0
  18. data/doc/user/Delegation.md +191 -0
  19. data/doc/user/GettingStarted.md +159 -0
  20. data/doc/user/ManagingContext.md +163 -0
  21. data/doc/user/MultiAgentWorkflows.md +256 -0
  22. data/doc/user/Python.md +159 -0
  23. data/doc/user/RunningInference.md +200 -0
  24. data/doc/user/ToolCalling.md +193 -0
  25. data/doc/user/WritingChats.md +197 -0
  26. data/lib/scout/llm/agent/chat.rb +61 -11
  27. data/lib/scout/llm/agent/delegate.rb +274 -65
  28. data/lib/scout/llm/agent/iterate.rb +2 -2
  29. data/lib/scout/llm/agent/save.rb +273 -0
  30. data/lib/scout/llm/agent/workflow.rb +164 -0
  31. data/lib/scout/llm/agent.rb +86 -61
  32. data/lib/scout/llm/ask.rb +62 -17
  33. data/lib/scout/llm/backends/anthropic.rb +9 -2
  34. data/lib/scout/llm/backends/bedrock.rb +15 -3
  35. data/lib/scout/llm/backends/default.rb +183 -99
  36. data/lib/scout/llm/backends/glm.rb +58 -0
  37. data/lib/scout/llm/backends/huggingface.rb +196 -26
  38. data/lib/scout/llm/backends/ollama.rb +13 -1
  39. data/lib/scout/llm/backends/openai.rb +0 -2
  40. data/lib/scout/llm/backends/openwebui.rb +20 -13
  41. data/lib/scout/llm/backends/relay.rb +22 -22
  42. data/lib/scout/llm/backends/responses.rb +1 -1
  43. data/lib/scout/llm/chat/agent_meta.rb +264 -0
  44. data/lib/scout/llm/chat/annotation.rb +39 -10
  45. data/lib/scout/llm/chat/parse.rb +28 -6
  46. data/lib/scout/llm/chat/persist.rb +25 -0
  47. data/lib/scout/llm/chat/process/clear.rb +41 -6
  48. data/lib/scout/llm/chat/process/files.rb +21 -6
  49. data/lib/scout/llm/chat/process/meta.rb +421 -34
  50. data/lib/scout/llm/chat/process/options.rb +21 -1
  51. data/lib/scout/llm/chat/process/tools.rb +56 -15
  52. data/lib/scout/llm/chat/process.rb +4 -0
  53. data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
  54. data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
  55. data/lib/scout/llm/chat/prompt.rb +48 -0
  56. data/lib/scout/llm/chat/provenance.rb +775 -0
  57. data/lib/scout/llm/chat/tool_calls.rb +76 -0
  58. data/lib/scout/llm/chat.rb +18 -2
  59. data/lib/scout/llm/embed.rb +11 -3
  60. data/lib/scout/llm/image.rb +86 -0
  61. data/lib/scout/llm/mcp.rb +10 -2
  62. data/lib/scout/llm/rag.rb +3 -3
  63. data/lib/scout/llm/tools/call.rb +160 -11
  64. data/lib/scout/llm/tools/knowledge_base.rb +1 -1
  65. data/lib/scout/llm/tools/workflow.rb +32 -16
  66. data/lib/scout/model/python/huggingface/causal.rb +23 -5
  67. data/lib/scout/model/python/huggingface.rb +2 -1
  68. data/lib/scout-ai.rb +1 -0
  69. data/python/README.md +197 -14
  70. data/python/scout_ai/huggingface/eval.py +245 -34
  71. data/python/tests/test_huggingface_eval.py +58 -0
  72. data/research/ChatAnalyst-required-changes.md +167 -0
  73. data/research/agent-delegation-analysis.md +810 -0
  74. data/research/agent-meta-provenance-integration-plan.md +622 -0
  75. data/research/agent-workflow-analysis.md +1120 -0
  76. data/research/backends-analysis.md +836 -0
  77. data/research/chat-core-analysis.md +946 -0
  78. data/research/chatanalyst-provenance/00-baseline.md +30 -0
  79. data/research/chatanalyst-provenance/01-repo-map.md +60 -0
  80. data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
  81. data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
  82. data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
  83. data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
  84. data/research/chatanalyst-provenance/07-critic-review.md +25 -0
  85. data/research/chatanalyst-provenance/final-report.md +45 -0
  86. data/research/chatanalyst-provenance/resumption.md +37 -0
  87. data/research/coding-philosophy-analysis.md +928 -0
  88. data/research/commands-analysis.md +947 -0
  89. data/research/multi-agent-patterns-analysis.md +853 -0
  90. data/research/prompt-strategies-analysis.md +630 -0
  91. data/research/prov-verbosity-fix-notes.md +77 -0
  92. data/research/provenance-analysis.md +469 -0
  93. data/research/provenance-navigation-design.md +640 -0
  94. data/research/synthesis-report.md +487 -0
  95. data/research/tools-system-analysis.md +779 -0
  96. data/scout-ai.gemspec +100 -11
  97. data/scout_commands/agent/ask +13 -3
  98. data/scout_commands/agent/kb +2 -0
  99. data/scout_commands/llm/ask +11 -4
  100. data/scout_commands/llm/md +76 -0
  101. data/scout_commands/llm/process_queries +48 -0
  102. data/scout_commands/llm/prov +602 -0
  103. data/scout_commands/llm/word +71 -0
  104. data/scout_commands/workflow/mcp +43 -0
  105. data/share/word/reference.docx +0 -0
  106. data/test/etc/AI/mock.yaml +11 -0
  107. data/test/fixtures/backends/anthropic.json +19 -0
  108. data/test/fixtures/backends/anthropic_tool_use.json +24 -0
  109. data/test/fixtures/backends/bedrock.json +8 -0
  110. data/test/fixtures/backends/bedrock_embedding.json +3 -0
  111. data/test/fixtures/backends/bedrock_tool_use.json +17 -0
  112. data/test/fixtures/backends/ollama.json +16 -0
  113. data/test/fixtures/backends/ollama_tool_call.json +27 -0
  114. data/test/fixtures/backends/openai_chat.json +21 -0
  115. data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
  116. data/test/fixtures/backends/responses.json +33 -0
  117. data/test/fixtures/backends/responses_tool_call.json +28 -0
  118. data/test/integration/README.md +32 -0
  119. data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
  120. data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
  121. data/test/integration/scout/llm/backends/test_relay.rb +52 -0
  122. data/test/integration/scout/llm/test_infrastructure.rb +74 -0
  123. data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
  124. data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
  125. data/test/integration/scout/model/test_base.rb +91 -0
  126. data/test/scout/llm/agent/test_chat.rb +8 -2
  127. data/test/scout/llm/agent/test_save.rb +413 -0
  128. data/test/scout/llm/agent/test_workflow.rb +110 -0
  129. data/test/scout/llm/backends/test_anthropic.rb +93 -10
  130. data/test/scout/llm/backends/test_bedrock.rb +118 -2
  131. data/test/scout/llm/backends/test_huggingface.rb +137 -42
  132. data/test/scout/llm/backends/test_ollama.rb +70 -20
  133. data/test/scout/llm/backends/test_openwebui.rb +42 -40
  134. data/test/scout/llm/backends/test_relay.rb +4 -2
  135. data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
  136. data/test/scout/llm/chat/process/test_meta.rb +518 -0
  137. data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
  138. data/test/scout/llm/chat/test_agent_meta.rb +357 -0
  139. data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
  140. data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
  141. data/test/scout/llm/chat/test_parse.rb +70 -15
  142. data/test/scout/llm/chat/test_prov_cli.rb +274 -0
  143. data/test/scout/llm/chat/test_provenance.rb +240 -0
  144. data/test/scout/llm/chat/test_tool_calls.rb +38 -0
  145. data/test/scout/llm/test_agent.rb +13 -36
  146. data/test/scout/llm/test_ask.rb +75 -52
  147. data/test/scout/llm/test_chat.rb +107 -13
  148. data/test/scout/llm/test_embed.rb +48 -0
  149. data/test/scout/llm/test_rag.rb +23 -16
  150. data/test/scout/llm/test_tools.rb +12 -1
  151. data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
  152. data/test/scout/llm/tools/test_mcp.rb +5 -3
  153. data/test/scout/llm/tools/test_workflow.rb +23 -2
  154. data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
  155. data/test/scout/model/python/huggingface/test_causal.rb +9 -3
  156. data/test/scout/model/python/huggingface/test_classification.rb +11 -2
  157. data/test/scout/model/python/test_torch.rb +2 -0
  158. data/test/scout/model/python/torch/test_helpers.rb +4 -0
  159. data/test/scout/model/test_base.rb +4 -2
  160. data/test/support/availability.rb +231 -0
  161. data/test/support/fake_clients.rb +138 -0
  162. data/test/support/fixtures.rb +21 -0
  163. data/test/support/infrastructure_probes.rb +136 -0
  164. data/test/support/mock_backend.rb +215 -0
  165. data/test/test_helper.rb +32 -2
  166. metadata +99 -10
  167. data/doc/Agent.md +0 -327
  168. data/doc/Chat.md +0 -458
  169. data/doc/LLM.md +0 -340
  170. data/doc/RAG.md +0 -129
  171. data/scout_commands/documenter +0 -148
  172. data/test/scout/llm/backends/test_openai.rb +0 -192
  173. data/test/scout/llm/backends/test_responses.rb +0 -238
  174. data/test/scout/llm/test_parse.rb +0 -98
@@ -0,0 +1,245 @@
1
+ # Design Principles
2
+
3
+ This document explains the coding philosophy and idioms that make Scout-AI
4
+ elegant and expressive. It is intended for framework contributors who want to
5
+ write code that fits the existing style.
6
+
7
+ > For detailed examples and analysis of idiomatic vs. non-idiomatic patterns,
8
+ > see [../../research/coding-philosophy-analysis.md](../../research/coding-philosophy-analysis.md).
9
+
10
+ ---
11
+
12
+ ## Abstraction-first
13
+
14
+ Every concept in Scout-AI is an abstraction with a crisp boundary:
15
+
16
+ - A **Chat** is "a conversation."
17
+ - An **Agent** is "a stateful conversation holder with tools."
18
+ - A **Backend** is "an adapter to a model API."
19
+ - A **Tool** is "a callable function exposed to the model."
20
+
21
+ New features should be expressed as a new abstraction or an extension of an
22
+ existing one, not as inline logic scattered across files. The question is
23
+ always: *what abstraction does this feature belong to?*
24
+
25
+ ---
26
+
27
+ ## Module composition over inheritance
28
+
29
+ Scout-AI avoids deep class hierarchies. Behavior is composed through Ruby
30
+ modules:
31
+
32
+ ```ruby
33
+ # Agent's behavior is split across multiple files that reopen the class:
34
+ # lib/scout/llm/agent.rb — core (ask, prompt, workflow)
35
+ # lib/scout/llm/agent/chat.rb — chat management
36
+ # lib/scout/llm/agent/iterate.rb — iteration patterns
37
+ # lib/scout/llm/agent/delegate.rb — multi-agent delegation
38
+ # lib/scout/llm/agent/workflow.rb — AgentWorkflow mixin + chat_task DSL
39
+ ```
40
+
41
+ Each module adds a cohesive set of methods to the same class. There is no
42
+ inheritance tree — just flat composition. This keeps each concern in its own
43
+ file while sharing `@other_options`, `@current_chat`, etc.
44
+
45
+ ---
46
+
47
+ ## Chat-as-data (annotate, don't wrap)
48
+
49
+ The single most important design decision:
50
+
51
+ > **A Chat is a plain `Array` of message `Hash`es, not an opaque object.**
52
+
53
+ The `Chat` module uses scout-essentials' `Annotation` system to add DSL methods
54
+ to a plain Array **non-invasively**:
55
+
56
+ ```ruby
57
+ chat = Chat.setup([])
58
+ chat.user("Hello")
59
+
60
+ chat.class # => Array (still an Array!)
61
+ chat.first[:role] # => "user"
62
+ chat.select { |m| m[:role] == 'system' } # standard Array operations work
63
+ ```
64
+
65
+ The annotation:
66
+ - Adds methods to the **singleton class** of the specific object instance.
67
+ - Does **not** change the object's class.
68
+ - Is **removable** via `Annotation.purge(obj)`.
69
+
70
+ **Implication:** Don't create wrapper classes for data that is already a Hash
71
+ or Array. Annotate it instead.
72
+
73
+ ---
74
+
75
+ ## Convention over configuration
76
+
77
+ Scout-AI discovers components by convention rather than registration:
78
+
79
+ | Convention | Resolution |
80
+ |---|---|
81
+ | Agent directory `Agent/<Name>/` | Auto-discovered via `Scout.Agent`, `Scout.chats.Agent`, etc. |
82
+ | `start_chat` file | Loaded as initial conversation. |
83
+ | `workflow.rb` | Loaded as the agent's workflow. |
84
+ | `knowledge_base/` | Loaded as the agent's KB. |
85
+ | `python/*.py` | Loaded as Python-backed tools. |
86
+
87
+ There are no registration calls or plugin manifests. Put files in the right
88
+ place and they are found.
89
+
90
+ ---
91
+
92
+ ## The DSL pattern
93
+
94
+ Scout-AI builds expressive domain-specific languages on top of Ruby's
95
+ flexibility:
96
+
97
+ ### Chat DSL
98
+
99
+ ```ruby
100
+ chat = Chat.setup([])
101
+ chat.system("You are a helpful assistant")
102
+ chat.user("What is 2+2?")
103
+ chat.option(:model, "gpt-4")
104
+ chat.ask # → "4"
105
+ ```
106
+
107
+ ### Agent DSL (via method_missing proxy)
108
+
109
+ ```ruby
110
+ agent = LLM.agent(model: "gpt-4")
111
+ agent.system("You are a coder")
112
+ agent.user("Write a function")
113
+ agent.chat
114
+ ```
115
+
116
+ The `method_missing` proxy on Agent forwards unknown methods to
117
+ `current_chat`, so all Chat methods work directly on the Agent.
118
+
119
+ ### Workflow DSL (chat_task)
120
+
121
+ ```ruby
122
+ module MyPipeline
123
+ extend Workflow
124
+ self.include_workflow AgentWorkflow
125
+
126
+ chat_task :my_task do
127
+ agent = self.agent :Worker, chat: chat
128
+ agent.user "Do something"
129
+ agent
130
+ end
131
+ end
132
+ ```
133
+
134
+ ---
135
+
136
+ ## Lazy initialization
137
+
138
+ Many things are initialized on first use, not eagerly:
139
+
140
+ ```ruby
141
+ def workflow(&block)
142
+ @workflow ||= begin
143
+ m = Module.new
144
+ m.extend Workflow
145
+ m.name ||= 'ScoutAgent'
146
+ m.tasks = {}
147
+ m
148
+ end
149
+ end
150
+
151
+ def current_chat
152
+ @current_chat ||= start
153
+ end
154
+ ```
155
+
156
+ This keeps object creation cheap and defers expensive setup until needed.
157
+
158
+ ---
159
+
160
+ ## IndiferentHash everywhere
161
+
162
+ Options and metadata use Scout's `IndiferentHash` (symbol/string-indifferent
163
+ access):
164
+
165
+ ```ruby
166
+ options[:model] # works
167
+ options['model'] # also works
168
+ ```
169
+
170
+ This eliminates a whole class of symbol-vs-string bugs. When building option
171
+ hashes, use `IndiferentHash.setup(hash)` rather than a plain Hash.
172
+
173
+ ---
174
+
175
+ ## Idiomatic patterns to follow
176
+
177
+ ### Use `Chat.setup` not `Chat.new`
178
+
179
+ ```ruby
180
+ # Good
181
+ chat = Chat.setup([])
182
+
183
+ # Wrong — Chat is a module, not a class
184
+ chat = Chat.new # NoMethodError
185
+ ```
186
+
187
+ ### Prefer annotation forwarding over explicit wrappers
188
+
189
+ ```ruby
190
+ # Good — Agent forwards to Chat via method_missing
191
+ agent.user("hello")
192
+
193
+ # Wrong — don't write explicit delegation methods
194
+ def agent_user(agent, msg)
195
+ agent.current_chat.user(msg)
196
+ end
197
+ ```
198
+
199
+ ### Use `Chat.follow` to compose conversations
200
+
201
+ ```ruby
202
+ # Good — follow prepends context
203
+ chat.follow(step(:plan).load)
204
+
205
+ # Avoid — manual concatenation loses annotations
206
+ chat = step(:plan).load + chat
207
+ ```
208
+
209
+ ### Prefer `proc` blocks for tools
210
+
211
+ ```ruby
212
+ # Good — Proc-based tool
213
+ LLM.add_tool(
214
+ "my_tool",
215
+ "Does something useful",
216
+ {"type" => "object", "properties" => {...}}
217
+ ) do |name, params|
218
+ # tool implementation
219
+ end
220
+ ```
221
+
222
+ ---
223
+
224
+ ## Anti-patterns to avoid
225
+
226
+ 1. **Creating wrapper classes for Arrays/Hashes** — Annotate instead.
227
+ 2. **Adding provider-specific logic to `LLM.ask`** — Put it in the Backend module.
228
+ 3. **Deep inheritance hierarchies** — Use module composition.
229
+ 4. **Eager initialization** — Use lazy `||=`.
230
+ 5. **Explicit delegation methods when `method_missing` already works** — Agent
231
+ already forwards to Chat; don't add `agent_user`, `agent_system`, etc.
232
+ 6. **Mutating the stored chat during prompt preparation** — Prompt strategies
233
+ are ephemeral; never mutate the source.
234
+ 7. **Using `Marshal.dump/load` for deep copying** — Procs can't be marshalled;
235
+ use `social_duplicate` patterns.
236
+ 8. **Hard-coding agent names in delegation logic** — Let the model choose via
237
+ `socialize`, or use `delegate` with explicit instances.
238
+
239
+ ---
240
+
241
+ ## Cross-references
242
+
243
+ - [Architecture.md](Architecture.md) — Overall system architecture.
244
+ - [ChatLifecycle.md](ChatLifecycle.md) — Chat data model.
245
+ - [../../research/coding-philosophy-analysis.md](../../research/coding-philosophy-analysis.md) — Deep investigation.
@@ -0,0 +1,292 @@
1
+ # Prompt Processing
2
+
3
+ This document explains the internal mechanism Scout-AI uses to manage long
4
+ contexts before sending a prompt to the LLM. It is intended for framework
5
+ contributors.
6
+
7
+ > For the user-facing guide on what happens when contexts get long, see
8
+ > [../user/ManagingContext.md](../user/ManagingContext.md).
9
+ > For deep code investigation, see
10
+ > [../../research/prompt-strategies-analysis.md](../../research/prompt-strategies-analysis.md).
11
+
12
+ ---
13
+
14
+ ## What prompt strategies are
15
+
16
+ Prompt strategies are a **pre-inference transformation layer** that modifies
17
+ the chat message list *just before* it is sent to the LLM backend. The primary
18
+ motivation is context-window management: in long agent conversations with many
19
+ tool calls, the accumulated arguments and return values can consume enormous
20
+ amounts of tokens.
21
+
22
+ The system works by applying named "strategies" to the message array. Each
23
+ strategy is a function that takes an Array of message hashes and returns a
24
+ (possibly shorter or modified) Array.
25
+
26
+ ---
27
+
28
+ ## File layout
29
+
30
+ Strategy implementations live in `lib/scout/llm/prompt/`:
31
+
32
+ ```
33
+ lib/scout/llm/
34
+ ├── chat/
35
+ │ └── prompt.rb # Dispatcher: prepare_prompt, shared constants
36
+ └── prompt/
37
+ ├── shorten_tools.rb # Default strategy (recomputes each turn)
38
+ └── shorten_tools_epoch.rb # Cache-friendly epoch variant
39
+ ```
40
+
41
+ The dispatcher in `prompt.rb` requires both strategy files and delegates to
42
+ them via a `case` statement inside `prepare_prompt`.
43
+
44
+ ---
45
+
46
+ ## The ephemeral design
47
+
48
+ **Prompt strategies never mutate the stored chat.** The transformation happens
49
+ entirely inside the backend's `ask` method:
50
+
51
+ ```ruby
52
+ # lib/scout/llm/backends/default.rb
53
+ prompt = Chat.prepare_prompt(messages, prompt_strategies)
54
+ ```
55
+
56
+ The variable `prompt` is a local derived from `messages`. The original messages
57
+ and the underlying Chat Object are untouched. This means:
58
+
59
+ - The chat history retains full-fidelity tool outputs for later inspection.
60
+ - The agent's persisted memory is not degraded by truncation.
61
+ - Only the *next inference* sees the shortened prompt.
62
+
63
+ This deliberately decouples *what the model sees* from *what the system
64
+ remembers*.
65
+
66
+ ---
67
+
68
+ ## `prepare_prompt` entry point
69
+
70
+ ```ruby
71
+ def self.prepare_prompt(prompt, prompt_strategies = nil)
72
+ ```
73
+
74
+ The method supports four input forms for `prompt_strategies`:
75
+
76
+ | Input type | Behavior |
77
+ |---|---|
78
+ | `Proc` | Called directly with the prompt array — full custom hook. |
79
+ | `nil` | Falls back to `DEFAULT_CONTEXT_STRATEGY` = `%w(shorten_tools)`. |
80
+ | `String` | Split by comma into strategy names (e.g., `"shorten_tools,custom"`). |
81
+ | `Array<String>` | Apply each named strategy in sequence. |
82
+
83
+ Strategies are applied **in sequence**: each receives the output of the previous.
84
+
85
+ The string `"none"` is a recognized no-op that returns the prompt unchanged.
86
+
87
+ ---
88
+
89
+ ## The `shorten_tools` strategy (default)
90
+
91
+ This is the **default strategy** — it runs on every backend `ask` call unless
92
+ explicitly disabled.
93
+
94
+ ### Algorithm
95
+
96
+ `shorten_tools` walks the message array **in reverse** (newest-first) and
97
+ applies a tiered truncation/dropping policy to `function_call` and
98
+ `function_call_output` messages. The most recent tool interactions get priority
99
+ for full retention; older ones are progressively degraded.
100
+
101
+ Three counters are tracked during the reverse traversal:
102
+
103
+ | Counter | Meaning |
104
+ |---|---|
105
+ | `tool_ids` | Count of tool output messages encountered (from the end) |
106
+ | `tool_chars` | Cumulative characters of retained tool content |
107
+ | `user_messages` | Number of user messages encountered |
108
+
109
+ ### Three-tier degradation
110
+
111
+ For tool outputs (the same logic applies to tool calls with separate thresholds):
112
+
113
+ | Position (from end) | Condition | Action |
114
+ |---|---|---|
115
+ | Most recent N | `count < full_tool_outputs` | **Full fidelity** |
116
+ | Middle band | Between `full_tool_outputs` and `max_tool_outputs` | **Truncated** (content shortened, hash-stamped) |
117
+ | Beyond max | `count > max_tool_outputs` | **Dropped entirely** |
118
+
119
+ There is also a character-budget override: if cumulative tool content is below
120
+ `max_tool_chars`, messages are kept at full fidelity regardless of position.
121
+ This means **short conversations are never truncated** — the system is a no-op
122
+ until context pressure is real.
123
+
124
+ ---
125
+
126
+ ## The `shorten_tools_epoch` strategy (cache-friendly)
127
+
128
+ ### Motivation
129
+
130
+ The default `shorten_tools` strategy recomputes the truncation boundary on
131
+ **every single inference**. When a new tool call is added, the boundary shifts
132
+ by one position, causing every previously-truncated message to be re-evaluated
133
+ with a different offset. This means the prompt prefix changes on every turn,
134
+ **defeating KV-cache and prompt-cache mechanisms** offered by LLM providers.
135
+
136
+ `shorten_tools_epoch` solves this by **freezing the compaction boundary** for
137
+ windows of N tool calls called *epochs*. Within an epoch, the compacted prefix
138
+ is byte-for-byte identical across consecutive inferences, maximizing cache hit
139
+ rates.
140
+
141
+ ### Algorithm
142
+
143
+ The conversation is divided into four regions (newest at the bottom):
144
+
145
+ ```
146
+ [ dropped ] tool calls older than (compacted + full) → removed entirely
147
+ [ compacted ] up to epoch_compacted_tool_calls tool calls, truncated
148
+ [ full-recent ] epoch_full_tool_calls tool calls at full fidelity
149
+ [ full-new ] any tool calls that arrived after the epoch boundary (full fidelity)
150
+ ```
151
+
152
+ The `compacted` and `full-recent` regions are **pinned** relative to the epoch
153
+ boundary, not the live tool-call count. Their content stays stable until the
154
+ boundary advances.
155
+
156
+ ### Epoch boundary calculation
157
+
158
+ ```
159
+ overflow = total_tool_calls - threshold # how many beyond threshold
160
+ epoch_idx = overflow > 0 ? (overflow - 1) / epoch_size : 0
161
+ pinned_total = threshold + (epoch_idx * epoch_size)
162
+ new_calls = total_tool_calls - pinned_total # tool calls that arrived this epoch
163
+ ```
164
+
165
+ The `pinned_total` determines where the `full-recent` region starts. Any tool
166
+ calls beyond `pinned_total` are treated as "new" and kept at full fidelity.
167
+
168
+ ### Worked example (threshold=100, full=10, compacted=40, epoch_size=10)
169
+
170
+ | Tool calls | pinned_total | new_calls | keep_full | compacted | dropped |
171
+ |---|---|---|---|---|---|
172
+ | 100 | 100 | 0 | 10 | 40 | 50 |
173
+ | 101 | 100 | 1 | 11 | 40 | 50 |
174
+ | 105 | 100 | 5 | 15 | 40 | 50 |
175
+ | 110 | 100 | 10 | 20 | 40 | 50 |
176
+ | 111 | 110 | 1 | 11 | 40 | 60 |
177
+
178
+ From tool calls 101–110 the compacted region (calls 11–50 from the pinned
179
+ boundary) is **identical**, so the prompt prefix is cache-stable for 10
180
+ consecutive inferences. At call 111 the boundary advances and the compacted
181
+ region shifts.
182
+
183
+ ### Configuration
184
+
185
+ | Config key | ENV var | Default | Description |
186
+ |---|---|---|---|
187
+ | `epoch_tool_call_threshold` | `EPOCH_TOOL_CALL_THRESHOLD` | 50 | Total tool calls at or below which no compaction happens |
188
+ | `epoch_full_tool_calls` | `EPOCH_FULL_TOOL_CALLS` | 10 | Most-recent tool calls kept at full fidelity |
189
+ | `epoch_compacted_tool_calls` | `EPOCH_COMPACTED_TOOL_CALLS` | 40 | Tool calls (before full-recent) to truncate |
190
+ | `epoch_size` | `EPOCH_SIZE` | 10 | New tool calls allowed before boundary advances |
191
+
192
+ All thresholds are read via `Scout::Config.get` and memoized in class variables,
193
+ following the same pattern as `shorten_tools`.
194
+
195
+ ### Enabling the epoch strategy
196
+
197
+ To use it instead of the default, pass the strategy name:
198
+
199
+ ```ruby
200
+ # In options
201
+ options[:prompt_strategies] = 'shorten_tools_epoch'
202
+
203
+ # Or directly
204
+ Chat.prepare_prompt(messages, 'shorten_tools_epoch')
205
+ ```
206
+
207
+ To switch the default system-wide, set:
208
+
209
+ ```ruby
210
+ Chat::DEFAULT_CONTEXT_STRATEGY.replace(['shorten_tools_epoch'])
211
+ ```
212
+
213
+ ---
214
+
215
+ ## Configuration thresholds (`shorten_tools`)
216
+
217
+ All thresholds are read via `Scout::Config.get` and **memoized** in class
218
+ variables on first access:
219
+
220
+ | Config key | ENV var | Default | Description |
221
+ |---|---|---|---|
222
+ | `full_tool_calls` | `FULL_TOOL_CALLS` | 0 | Recent tool calls kept at full fidelity |
223
+ | `full_tool_outputs` | `FULL_TOOL_OUTPUTS` | 10 | Recent tool outputs kept at full fidelity |
224
+ | `max_tool_calls` | `MAX_TOOL_CALLS` | 40 | Hard limit; tool calls beyond this are dropped |
225
+ | `max_tool_outputs` | `MAX_TOOL_OUTPUTS` | 40 (defaults to `max_tool_calls`) | Hard limit; outputs beyond this are dropped |
226
+ | `max_tool_chars` | `MAX_TOOL_CHARS` | 100,000 | Cumulative character budget for retained tool content |
227
+
228
+ ### Memoization trade-off
229
+
230
+ Because thresholds use `||=` memoization, they are **frozen for the process
231
+ lifetime** after first access. Changing config files or ENV vars mid-process
232
+ has no effect. This is fine for CLI/agent usage but could be surprising in
233
+ long-running daemons.
234
+
235
+ ---
236
+
237
+ ## Content hashing in truncated strings
238
+
239
+ When a value is truncated, `Log.truncate_string` embeds an MD5 hash prefix:
240
+
241
+ ```
242
+ Truncated (15432): The first ~70 chars...<...15432 - a1b2c...>...last ~70 chars
243
+ ```
244
+
245
+ This allows truncated content to be matched against logs or the original chat
246
+ for debugging.
247
+
248
+ ---
249
+
250
+ ## Integration with the backend
251
+
252
+ `prepare_prompt` is called inside `Backend::Default#ask`, in the normal
253
+ (non-relay) path:
254
+
255
+ ```ruby
256
+ client = prepare_client(options, messages)
257
+ prompt = Chat.prepare_prompt(messages, prompt_strategies)
258
+ formatted_prompt = format_messages(prompt)
259
+ tools = tools(formatted_prompt, options)
260
+ response = query(client, formatted_prompt, tools, options)
261
+ ```
262
+
263
+ Key points:
264
+ - Applied on every `ask` invocation in the normal path.
265
+ - **Not applied in relay mode** (raw messages are uploaded to a remote server).
266
+ - Applied **before** `format_messages` — strategy output directly determines
267
+ token consumption.
268
+ - `prompt_strategies` comes from the `options` hash, so callers can override
269
+ per-call.
270
+
271
+ ---
272
+
273
+ ## Extension point: custom strategies
274
+
275
+ Two mechanisms coexist:
276
+
277
+ 1. **Hard-coded `case` dispatch** for built-in strategies (`shorten_tools`,
278
+ `shorten_tools_epoch`, `none`).
279
+ 2. **`REGISTERED_STRATEGIES` hash** for user/plugin-registered strategies.
280
+
281
+ > **Note:** `REGISTERED_STRATEGIES` is referenced in the code but not yet
282
+ > populated with entries. Passing an unknown strategy name will result in
283
+ > `nil.call(prompt)`, raising a `NoMethodError`. For now, use a `Proc` to
284
+ > supply custom strategies.
285
+
286
+ ---
287
+
288
+ ## Cross-references
289
+
290
+ - [../user/ManagingContext.md](../user/ManagingContext.md) — User guide for long contexts.
291
+ - [Backends.md](Backends.md) — Where `prepare_prompt` is called in the inference loop.
292
+ - [../../research/prompt-strategies-analysis.md](../../research/prompt-strategies-analysis.md) — Deep investigation.