scout-ai 1.2.3 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (174) hide show
  1. checksums.yaml +4 -4
  2. data/.vimproject +138 -50
  3. data/README.md +171 -290
  4. data/Rakefile +17 -1
  5. data/VERSION +1 -1
  6. data/doc/Improvements.md +325 -0
  7. data/doc/StartHere.md +110 -0
  8. data/doc/developer/Architecture.md +126 -0
  9. data/doc/developer/Backends.md +199 -0
  10. data/doc/developer/ChatLifecycle.md +183 -0
  11. data/doc/developer/DelegationInternals.md +295 -0
  12. data/doc/developer/DesignPrinciples.md +245 -0
  13. data/doc/developer/PromptProcessing.md +292 -0
  14. data/doc/developer/Provenance.md +317 -0
  15. data/doc/user/BuildingAgents.md +345 -0
  16. data/doc/user/Cookbook.md +333 -0
  17. data/doc/user/CoreConcepts.md +181 -0
  18. data/doc/user/Delegation.md +191 -0
  19. data/doc/user/GettingStarted.md +159 -0
  20. data/doc/user/ManagingContext.md +163 -0
  21. data/doc/user/MultiAgentWorkflows.md +256 -0
  22. data/doc/user/Python.md +159 -0
  23. data/doc/user/RunningInference.md +200 -0
  24. data/doc/user/ToolCalling.md +193 -0
  25. data/doc/user/WritingChats.md +197 -0
  26. data/lib/scout/llm/agent/chat.rb +61 -11
  27. data/lib/scout/llm/agent/delegate.rb +274 -65
  28. data/lib/scout/llm/agent/iterate.rb +2 -2
  29. data/lib/scout/llm/agent/save.rb +273 -0
  30. data/lib/scout/llm/agent/workflow.rb +164 -0
  31. data/lib/scout/llm/agent.rb +86 -61
  32. data/lib/scout/llm/ask.rb +62 -17
  33. data/lib/scout/llm/backends/anthropic.rb +9 -2
  34. data/lib/scout/llm/backends/bedrock.rb +15 -3
  35. data/lib/scout/llm/backends/default.rb +183 -99
  36. data/lib/scout/llm/backends/glm.rb +58 -0
  37. data/lib/scout/llm/backends/huggingface.rb +196 -26
  38. data/lib/scout/llm/backends/ollama.rb +13 -1
  39. data/lib/scout/llm/backends/openai.rb +0 -2
  40. data/lib/scout/llm/backends/openwebui.rb +20 -13
  41. data/lib/scout/llm/backends/relay.rb +22 -22
  42. data/lib/scout/llm/backends/responses.rb +1 -1
  43. data/lib/scout/llm/chat/agent_meta.rb +264 -0
  44. data/lib/scout/llm/chat/annotation.rb +39 -10
  45. data/lib/scout/llm/chat/parse.rb +28 -6
  46. data/lib/scout/llm/chat/persist.rb +25 -0
  47. data/lib/scout/llm/chat/process/clear.rb +41 -6
  48. data/lib/scout/llm/chat/process/files.rb +21 -6
  49. data/lib/scout/llm/chat/process/meta.rb +421 -34
  50. data/lib/scout/llm/chat/process/options.rb +21 -1
  51. data/lib/scout/llm/chat/process/tools.rb +56 -15
  52. data/lib/scout/llm/chat/process.rb +4 -0
  53. data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
  54. data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
  55. data/lib/scout/llm/chat/prompt.rb +48 -0
  56. data/lib/scout/llm/chat/provenance.rb +775 -0
  57. data/lib/scout/llm/chat/tool_calls.rb +76 -0
  58. data/lib/scout/llm/chat.rb +18 -2
  59. data/lib/scout/llm/embed.rb +11 -3
  60. data/lib/scout/llm/image.rb +86 -0
  61. data/lib/scout/llm/mcp.rb +10 -2
  62. data/lib/scout/llm/rag.rb +3 -3
  63. data/lib/scout/llm/tools/call.rb +160 -11
  64. data/lib/scout/llm/tools/knowledge_base.rb +1 -1
  65. data/lib/scout/llm/tools/workflow.rb +32 -16
  66. data/lib/scout/model/python/huggingface/causal.rb +23 -5
  67. data/lib/scout/model/python/huggingface.rb +2 -1
  68. data/lib/scout-ai.rb +1 -0
  69. data/python/README.md +197 -14
  70. data/python/scout_ai/huggingface/eval.py +245 -34
  71. data/python/tests/test_huggingface_eval.py +58 -0
  72. data/research/ChatAnalyst-required-changes.md +167 -0
  73. data/research/agent-delegation-analysis.md +810 -0
  74. data/research/agent-meta-provenance-integration-plan.md +622 -0
  75. data/research/agent-workflow-analysis.md +1120 -0
  76. data/research/backends-analysis.md +836 -0
  77. data/research/chat-core-analysis.md +946 -0
  78. data/research/chatanalyst-provenance/00-baseline.md +30 -0
  79. data/research/chatanalyst-provenance/01-repo-map.md +60 -0
  80. data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
  81. data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
  82. data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
  83. data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
  84. data/research/chatanalyst-provenance/07-critic-review.md +25 -0
  85. data/research/chatanalyst-provenance/final-report.md +45 -0
  86. data/research/chatanalyst-provenance/resumption.md +37 -0
  87. data/research/coding-philosophy-analysis.md +928 -0
  88. data/research/commands-analysis.md +947 -0
  89. data/research/multi-agent-patterns-analysis.md +853 -0
  90. data/research/prompt-strategies-analysis.md +630 -0
  91. data/research/prov-verbosity-fix-notes.md +77 -0
  92. data/research/provenance-analysis.md +469 -0
  93. data/research/provenance-navigation-design.md +640 -0
  94. data/research/synthesis-report.md +487 -0
  95. data/research/tools-system-analysis.md +779 -0
  96. data/scout-ai.gemspec +100 -11
  97. data/scout_commands/agent/ask +13 -3
  98. data/scout_commands/agent/kb +2 -0
  99. data/scout_commands/llm/ask +11 -4
  100. data/scout_commands/llm/md +76 -0
  101. data/scout_commands/llm/process_queries +48 -0
  102. data/scout_commands/llm/prov +602 -0
  103. data/scout_commands/llm/word +71 -0
  104. data/scout_commands/workflow/mcp +43 -0
  105. data/share/word/reference.docx +0 -0
  106. data/test/etc/AI/mock.yaml +11 -0
  107. data/test/fixtures/backends/anthropic.json +19 -0
  108. data/test/fixtures/backends/anthropic_tool_use.json +24 -0
  109. data/test/fixtures/backends/bedrock.json +8 -0
  110. data/test/fixtures/backends/bedrock_embedding.json +3 -0
  111. data/test/fixtures/backends/bedrock_tool_use.json +17 -0
  112. data/test/fixtures/backends/ollama.json +16 -0
  113. data/test/fixtures/backends/ollama_tool_call.json +27 -0
  114. data/test/fixtures/backends/openai_chat.json +21 -0
  115. data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
  116. data/test/fixtures/backends/responses.json +33 -0
  117. data/test/fixtures/backends/responses_tool_call.json +28 -0
  118. data/test/integration/README.md +32 -0
  119. data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
  120. data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
  121. data/test/integration/scout/llm/backends/test_relay.rb +52 -0
  122. data/test/integration/scout/llm/test_infrastructure.rb +74 -0
  123. data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
  124. data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
  125. data/test/integration/scout/model/test_base.rb +91 -0
  126. data/test/scout/llm/agent/test_chat.rb +8 -2
  127. data/test/scout/llm/agent/test_save.rb +413 -0
  128. data/test/scout/llm/agent/test_workflow.rb +110 -0
  129. data/test/scout/llm/backends/test_anthropic.rb +93 -10
  130. data/test/scout/llm/backends/test_bedrock.rb +118 -2
  131. data/test/scout/llm/backends/test_huggingface.rb +137 -42
  132. data/test/scout/llm/backends/test_ollama.rb +70 -20
  133. data/test/scout/llm/backends/test_openwebui.rb +42 -40
  134. data/test/scout/llm/backends/test_relay.rb +4 -2
  135. data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
  136. data/test/scout/llm/chat/process/test_meta.rb +518 -0
  137. data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
  138. data/test/scout/llm/chat/test_agent_meta.rb +357 -0
  139. data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
  140. data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
  141. data/test/scout/llm/chat/test_parse.rb +70 -15
  142. data/test/scout/llm/chat/test_prov_cli.rb +274 -0
  143. data/test/scout/llm/chat/test_provenance.rb +240 -0
  144. data/test/scout/llm/chat/test_tool_calls.rb +38 -0
  145. data/test/scout/llm/test_agent.rb +13 -36
  146. data/test/scout/llm/test_ask.rb +75 -52
  147. data/test/scout/llm/test_chat.rb +107 -13
  148. data/test/scout/llm/test_embed.rb +48 -0
  149. data/test/scout/llm/test_rag.rb +23 -16
  150. data/test/scout/llm/test_tools.rb +12 -1
  151. data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
  152. data/test/scout/llm/tools/test_mcp.rb +5 -3
  153. data/test/scout/llm/tools/test_workflow.rb +23 -2
  154. data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
  155. data/test/scout/model/python/huggingface/test_causal.rb +9 -3
  156. data/test/scout/model/python/huggingface/test_classification.rb +11 -2
  157. data/test/scout/model/python/test_torch.rb +2 -0
  158. data/test/scout/model/python/torch/test_helpers.rb +4 -0
  159. data/test/scout/model/test_base.rb +4 -2
  160. data/test/support/availability.rb +231 -0
  161. data/test/support/fake_clients.rb +138 -0
  162. data/test/support/fixtures.rb +21 -0
  163. data/test/support/infrastructure_probes.rb +136 -0
  164. data/test/support/mock_backend.rb +215 -0
  165. data/test/test_helper.rb +32 -2
  166. metadata +99 -10
  167. data/doc/Agent.md +0 -327
  168. data/doc/Chat.md +0 -458
  169. data/doc/LLM.md +0 -340
  170. data/doc/RAG.md +0 -129
  171. data/scout_commands/documenter +0 -148
  172. data/test/scout/llm/backends/test_openai.rb +0 -192
  173. data/test/scout/llm/backends/test_responses.rb +0 -238
  174. data/test/scout/llm/test_parse.rb +0 -98
@@ -0,0 +1,640 @@
1
+ # Provenance navigation in Scout-AI
2
+
3
+ > This is an investigation and design proposal, not normative documentation.
4
+ > It reflects the code inspected in `scout-ai` and the current
5
+ > `/bulk/mvazque2/git/workflows/ChatAnalyst` checkout. It intentionally does
6
+ > not propose implementations for agent monitoring, `ask_user`, or explicit
7
+ > agent-raised failure; those topics are considered only where they constrain
8
+ > provenance.
9
+
10
+ ## Executive conclusion
11
+
12
+ The difficult part of Scout-AI provenance is not recursion. It is that a run
13
+ interleaves two already-valid Scout data models:
14
+
15
+ 1. a Chat is an Array of messages and contains conversational lineage,
16
+ inference metadata, imports, and job projections;
17
+ 2. a Step is a persisted Workflow execution and contains dependencies, status,
18
+ inputs, result files, and log artifacts.
19
+
20
+ The provenance structure is therefore a heterogeneous directed graph whose
21
+ only persistent node kinds are **chat files** and **workflow jobs**. Agents,
22
+ tool calls, inference segments, and token events are observations inside those
23
+ nodes; they should not be promoted into Session, ProvenanceContext, ChatGraph,
24
+ or Node classes merely to make traversal convenient.
25
+
26
+ The core library should provide one small, block-oriented traversal primitive
27
+ that walks native `Path` and `Step` values and yields typed relationships. The
28
+ `prov` command and ChatAnalyst should consume that primitive and independently
29
+ format or analyse what they receive. Message tracing, tool-call pairing, and
30
+ token accounting should remain separate Chat operations layered on top of the
31
+ files returned by traversal.
32
+
33
+ The most important data-model gap is more fundamental: direct inference meta
34
+ messages currently have no persisted unique inference/event identifier. The
35
+ lineage digest is excellent for recognizing copied conversation history, but
36
+ it cannot distinguish two genuinely repeated requests with identical history
37
+ and metadata. Consequently, current token deduplication is useful but cannot
38
+ be exact in every case. A locally generated inference ID should eventually be
39
+ persisted in each direct meta record. This is independent of traversal and
40
+ should not block the traversal cleanup.
41
+
42
+ ## Sources examined
43
+
44
+ The investigation covered:
45
+
46
+ - `scout_commands/llm/prov`
47
+ - `lib/scout/llm/chat/process/meta.rb`
48
+ - `lib/scout/llm/chat/provenance.rb`
49
+ - `lib/scout/llm/agent/workflow.rb`
50
+ - `lib/scout/llm/agent/delegate.rb`
51
+ - `lib/scout/llm/tools/call.rb`
52
+ - `lib/scout/llm/backends/default.rb`
53
+ - provenance and meta tests under `test/scout/llm/`
54
+ - the current and previous ChatAnalyst implementations in
55
+ `/bulk/mvazque2/git/workflows/ChatAnalyst/workflow.rb` and
56
+ `workflow.rb.orig`
57
+ - persisted jobs and recursively nested agent logs under
58
+ `~/.scout/var/jobs`
59
+ - agent examples under `~/chats/Agent`
60
+ - Scout Workflow lifecycle and provenance documentation
61
+
62
+ A real `Refined/ask` result with separate
63
+ `log/worker-round-1/agent.chat` and `log/critic-round-1/agent.chat` files was
64
+ used as a concrete layout check. The current ChatAnalyst Session found five
65
+ chats, two jobs, and six edges from one sampled result, which confirms that the
66
+ basic recursive idea works. It does not remove the design and correctness
67
+ issues below.
68
+
69
+ ## The crisp model
70
+
71
+ ### Persistent facts
72
+
73
+ A provenance reader should begin with facts that are already persisted, not
74
+ with inferred conceptual objects.
75
+
76
+ A chat file can directly state:
77
+
78
+ - conversational messages;
79
+ - `import`, `continue`, and `last` references to other chat files;
80
+ - direct inference metadata in `meta` messages;
81
+ - a producer job reference in `meta job=...`;
82
+ - function calls and function-call outputs.
83
+
84
+ A workflow job can directly state or expose:
85
+
86
+ - its path and logical Workflow/task identity;
87
+ - status, timing, exception, inputs, and dependency paths in Step info;
88
+ - its result, which may itself be a chat file;
89
+ - recursively named chat artifacts under `.files/log/**/*.chat`.
90
+
91
+ Nothing else is required to navigate the structural provenance.
92
+
93
+ ### Structural relations
94
+
95
+ Traversal from a requested root follows five relations:
96
+
97
+ | Parent | Relation | Child | Meaning |
98
+ |---|---|---|---|
99
+ | chat | `import` | chat | The parent incorporates another persisted chat. |
100
+ | chat | `job` | job | A projected response in the chat was produced by the job. |
101
+ | job | `dependency` | job | The job consumes a normal Scout dependency. |
102
+ | job | `log` | chat | The job persisted an agent conversation as a log artifact. |
103
+ | job | `result` | chat | The job result itself is a chat and must be inspected, especially when traversal starts from a Step. |
104
+
105
+ These are **root-outward discovery relations**. They are deliberately not
106
+ presentation arrows. For example, a flow diagram may render a producer job as
107
+ feeding a chat, while recursive discovery follows the chat's `job` reference
108
+ toward the producer. `prov` currently mixes discovery hierarchy with natural
109
+ data-flow direction, which makes otherwise simple code hard to reason about.
110
+ Renderers should decide arrow direction after traversal.
111
+
112
+ ### Observations inside nodes
113
+
114
+ The following are analyses of node content, not additional structural node
115
+ kinds:
116
+
117
+ - an inference segment is a meta record plus the messages it covers;
118
+ - a tool invocation is a paired `function_call` and
119
+ `function_call_output`;
120
+ - agent delegation is a tool invocation whose tool is `ask` or
121
+ `hand_off_to_*`;
122
+ - an agent log is still a chat file; “agent” is a label inferred from its log
123
+ path or explicit future metadata;
124
+ - token usage is attached to direct inference segments;
125
+ - a job failure or abort is Step status and exception evidence.
126
+
127
+ Keeping this distinction prevents the graph from becoming a second runtime
128
+ object model.
129
+
130
+ ## What currently works well
131
+
132
+ Several core decisions are strong and should be retained.
133
+
134
+ ### Chat remains plain data
135
+
136
+ `Chat.load` reads persisted chat text without recompiling roles. This is the
137
+ right safety boundary for provenance: inspection must never execute imports,
138
+ tasks, files, or tools.
139
+
140
+ ### Job projections are explicit
141
+
142
+ `Chat.project` removes direct meta records from the projected output and adds a
143
+ single `meta job=<path>` marker. This cleanly distinguishes “tokens spent in
144
+ this chat” from “content produced elsewhere and projected here.”
145
+
146
+ ### Conversational lineage is content-addressed
147
+
148
+ `message_index` hashes each non-meta message with its conversational
149
+ predecessor. Copied or inherited histories therefore share lineage identities,
150
+ while changed histories diverge. Excluding meta from the conversational chain
151
+ is correct because meta is not provider input.
152
+
153
+ ### Segment tracing is independent of job traversal
154
+
155
+ `trace_indices`, `trace_chats`, `direct_entries`, and `token_totals` operate on
156
+ Chats and do not need to know how files were found. This separation is exactly
157
+ the direction the rest of the provenance implementation should follow.
158
+
159
+ ### Workflow provenance is already authoritative for jobs
160
+
161
+ Step info already records dependencies, lifecycle status, timings, exceptions,
162
+ and other execution facts. Scout-AI should link to that evidence rather than
163
+ replicate it in a Session-like object.
164
+
165
+ ## Current problems
166
+
167
+ ### Traversal is duplicated three times
168
+
169
+ The same heterogeneous walk is independently implemented in:
170
+
171
+ - `scout_commands/llm/prov#report`;
172
+ - `Chat.provenance` and related helpers;
173
+ - `ChatAnalyst::Session` (and previously the much larger
174
+ `ChatAnalyst::ChatGraph`).
175
+
176
+ Each copy makes slightly different decisions about imports, result chats,
177
+ dependencies, logs, deduplication, errors, and edge direction. Fixes therefore
178
+ do not propagate to all consumers.
179
+
180
+ ### Core helper contracts are inconsistent
181
+
182
+ Current examples include:
183
+
184
+ - `Chat.job_chat_files(job)` recursively includes dependencies;
185
+ - `Chat.job_agent_chat_files(job)` includes only the given job's logs;
186
+ - the instance `job_agent_chat_files` only expands the chat's immediate jobs;
187
+ - documentation describes some of these as recursively traversing all
188
+ dependencies;
189
+ - `job_agent_chats` is named as if it returned loaded Chat values but currently
190
+ returns paths.
191
+
192
+ A caller cannot infer recursion or return type from these names reliably.
193
+ Traversal should be centralized, and direct-neighbour helpers should say that
194
+ they are direct.
195
+
196
+ ### `Chat.provenance` is not a faithful general traversal
197
+
198
+ The current implementation follows job logs and then calls
199
+ `Chat.provenance(dep.path)` for dependencies. That treats an arbitrary Step
200
+ result path as if it were necessarily a chat. It also omits chat imports and
201
+ uses a nested Hash whose meaning is limited to chat-to-log relationships;
202
+ normal dependency and producer relationships are not represented explicitly.
203
+ Broad existence checks and implicit rescues obscure malformed or missing
204
+ evidence.
205
+
206
+ ### `prov` combines discovery, printing, graph construction, token analysis,
207
+ filtering, naming, and rendering
208
+
209
+ The recursive `report` both prints and builds a nested graph. Later flow code
210
+ must reverse or reinterpret edges based on whether a Hash key happens to be a
211
+ Step or a String. Global instance variables cache nodes, edges, chats, and
212
+ tokens. This makes the command longer and less reliable than the underlying
213
+ problem warrants.
214
+
215
+ There are also concrete fragilities:
216
+
217
+ - identity alternates between Step objects and paths;
218
+ - `report` may return `nil` for a seen object while callers expect a Hash to
219
+ merge;
220
+ - imports are absent;
221
+ - an unused `load_chat` helper remains;
222
+ - root detection relies on the presence of `<filename>.files`;
223
+ - display suppression of `agent.chat` is entangled with token attribution;
224
+ - “agent” type is guessed from `'.files/log/'` in a path.
225
+
226
+ ### ChatAnalyst's Session is a cache plus recursive side effects
227
+
228
+ `Session` eagerly resolves and traverses everything in its constructor, stores
229
+ four mutable collections, embeds path resolution and warning policy, and adds
230
+ token methods that already exist on Chat. Every task constructs a fresh
231
+ Session and then reads those caches.
232
+
233
+ This abstraction adds no domain concept. It is an execution context for one
234
+ algorithm. More importantly, it encourages future agent-written code to add
235
+ more convenience methods until traversal, accounting, tool semantics, and
236
+ reporting are coupled again.
237
+
238
+ The earlier `ChatGraph` demonstrates that failure mode clearly: it grew path
239
+ canonicalization, hidden-path policy, tool parsing, success inference, usage
240
+ accounting, overlap analysis, message storage, job fallbacks, and graph
241
+ construction inside one class. Replacing it with a smaller Session reduced the
242
+ amount of code but not the architectural cause.
243
+
244
+ ### Error handling is mostly invisible
245
+
246
+ Some library traversal helpers rescue all exceptions and return empty arrays.
247
+ ChatAnalyst catches exceptions and stores warning strings. `prov` often lets
248
+ errors escape. These policies make the same missing log or unreadable legacy
249
+ Step look like “no provenance,” a warning, or a fatal error depending on the
250
+ consumer.
251
+
252
+ The traversal primitive should not invent an error-monitoring subsystem, but
253
+ it must expose failures with their source node and attempted relation so that a
254
+ CLI can warn, an analyst can report incomplete evidence, and a strict caller
255
+ can raise.
256
+
257
+ ### Location identity and lineage identity are conflated
258
+
259
+ There are two legitimate identities:
260
+
261
+ - a **message address**, such as `[chat_path, index]`, identifies where a
262
+ persisted message can be retrieved;
263
+ - a **lineage ID** identifies equivalent conversational content and is useful
264
+ for recognizing inherited or copied history.
265
+
266
+ ChatAnalyst creates strings such as `path#index` and calls them IDs, while Chat
267
+ also calls the lineage digest an ID. Reports should name these fields
268
+ `address` and `lineage_id` explicitly. An address should remain structured
269
+ until final JSON formatting rather than relying on parsing a path containing a
270
+ separator.
271
+
272
+ ### Exact inference identity is not persisted
273
+
274
+ `Backend::Default#update_meta` currently stores token fields and cumulative
275
+ snapshots, but the tests explicitly assert that `usage_id` is absent.
276
+ `trace_chats` deduplicates segments by the lineage-derived meta ID. This is
277
+ correct for a copied historical segment but can collapse two genuinely
278
+ repeated requests when all of the following are identical:
279
+
280
+ - preceding conversational lineage;
281
+ - meta token fields;
282
+ - response content.
283
+
284
+ Conversely, session counters are process snapshots and cannot identify a
285
+ request. Exact cost and execution counting requires a unique persisted direct
286
+ inference identity generated by Scout-AI, independent of provider request IDs.
287
+
288
+ ### Delegation is not always structurally linked to its log
289
+
290
+ Workflow-backed agent calls have strong job and log links. Socialized agents
291
+ are returned by the `ask` tool, executed later by `process_calls`, and finally
292
+ written by `log_agent` under a society path. The function call says which
293
+ agent/conversation was requested, and the path convention often allows a
294
+ match, but no explicit durable ID links that call to the precise specialist
295
+ chat or inference segment. This should be reported as a semantic association,
296
+ not presented as an authoritative structural edge.
297
+
298
+ ## Proposed core primitives
299
+
300
+ ### 1. Direct-neighbour readers
301
+
302
+ First make direct facts explicit and non-recursive. Suggested responsibilities
303
+ are:
304
+
305
+ - resolve a persisted chat reference without compiling it;
306
+ - return direct import files for a chat and its source path;
307
+ - return direct job references from a chat;
308
+ - return direct dependencies of a Step;
309
+ - return direct log chat files of a Step;
310
+ - return the Step result chat when its type is chat.
311
+
312
+ Existing APIs can supply several of these, but names and contracts should be
313
+ made consistent. In particular, methods named `*_chats` should return Chat
314
+ values and methods named `*_files` should return paths. Recursion should not be
315
+ hidden inside either.
316
+
317
+ ### 2. One block-oriented traversal
318
+
319
+ Add a module function on Chat, not a class. A possible contract is:
320
+
321
+ Chat.traverse_provenance(root, follow: :all, on_error: nil) do |
322
+ kind, object, parent_kind, parent, relation
323
+ |
324
+ # kind is :chat or :job
325
+ # object is a Path for :chat and a Step for :job
326
+ # root has nil parent and relation
327
+ end
328
+
329
+ Without a block it should return an Enumerator. The implementation needs only
330
+ a queue or recursion and a Set. The visited key must include node kind, because
331
+ a chat-typed job result may have the same filesystem path as its Step.
332
+
333
+ The traversal should:
334
+
335
+ 1. normalize a root explicitly as a chat file or Step;
336
+ 2. yield the root once;
337
+ 3. read only direct neighbours;
338
+ 4. yield each newly visited child together with the relation that discovered
339
+ it;
340
+ 5. retain cycle safety and shared dependency deduplication;
341
+ 6. preserve native values rather than wrapping them in node classes;
342
+ 7. expose read/load failures through `on_error` with the parent and relation.
343
+
344
+ The exact block argument order can change, but the essential contract is that
345
+ consumers receive native objects plus explicit type and relation. A flat stream
346
+ is sufficient to print a tree, collect JSON nodes and edges, calculate tokens,
347
+ or produce DOT.
348
+
349
+ An error callback can receive:
350
+
351
+ error, kind, object, relation, referenced_value
352
+
353
+ If no callback is supplied, raising is the clearest library default. CLI and
354
+ analysis callers can opt into “record warning and continue.” Silent rescue and
355
+ empty output should not be the default because absence of evidence differs
356
+ from unreadable evidence.
357
+
358
+ ### 3. Small collectors implemented from traversal
359
+
360
+ Convenience is still useful when it preserves the same semantics. Small module
361
+ functions can collect from the stream:
362
+
363
+ - `Chat.provenance_chat_files(root)`;
364
+ - `Chat.provenance_jobs(root)`;
365
+ - `Chat.provenance_edges(root)`.
366
+
367
+ These should be thin collectors, not alternate traversal implementations.
368
+ The existing `Chat.provenance` can either become a compatibility collector or
369
+ be deprecated after consumers migrate.
370
+
371
+ ### 4. Source-aware tracing
372
+
373
+ Keep `Chat.trace_chats` for compatibility, but add a source-aware form that
374
+ accepts `[path, chat]` pairs and retains addresses:
375
+
376
+ { meta_address: [path, index],
377
+ lineage_id: ...,
378
+ meta: ...,
379
+ messages: [[path, index], ...],
380
+ message_lineages: [...] }
381
+
382
+ Deduplication should remain based on explicit inference ID when present and
383
+ fall back to the current lineage rule for legacy chats. Retrieval should use
384
+ addresses; overlap analysis should use lineage IDs.
385
+
386
+ ### 5. Tool-call pairing as a Chat primitive
387
+
388
+ Pairing calls and outputs is message analysis used by any future analyst, not
389
+ only ChatAnalyst. Add a small operation that returns plain Hash records and
390
+ preserves both addresses. It should support current provider-normalized roles
391
+ and fields (`function_call`, `mcp_call`, `function_call_output`, `id`,
392
+ `call_id`, nested function name).
393
+
394
+ Success/failure interpretation should be separate. Pairing can authoritatively
395
+ say whether an output exists and what it contains. Deciding that JSON with
396
+ `exit_status != 0` means a failed shell call is a reporting policy, not generic
397
+ Chat structure. A helper may provide the common policy, but the raw pair must
398
+ remain available.
399
+
400
+ ## Suggested traversal implementation shape
401
+
402
+ The implementation can be short and procedural:
403
+
404
+ 1. a queue contains tuples of kind, native object, parent kind, parent, and
405
+ relation;
406
+ 2. a Set contains `[kind, canonical_path]` keys;
407
+ 3. visiting a chat loads it once, enqueues direct imports and producer jobs;
408
+ 4. visiting a job enqueues direct dependencies, logs, and its result chat;
409
+ 5. each loading operation is wrapped only at the boundary needed to call the
410
+ configured error handler.
411
+
412
+ There is no need for Session state, a Graph object, node subclasses, edge
413
+ classes, visitor classes, or a provenance database.
414
+
415
+ A consumer that needs a graph can use ordinary Hashes and Arrays:
416
+
417
+ nodes = {}
418
+ edges = []
419
+
420
+ Chat.traverse_provenance(root, on_error: record_warning) do |
421
+ kind, object, parent_kind, parent, relation
422
+ |
423
+ key = [kind, provenance_path(object)]
424
+ nodes[key] ||= object
425
+ edges << [[parent_kind, provenance_path(parent)], relation, key] if parent
426
+ end
427
+
428
+ That state belongs to the report being built, not to the core traversal.
429
+
430
+ ## How `prov` should change
431
+
432
+ `prov` should become three clearly separated stages.
433
+
434
+ ### Discovery
435
+
436
+ Call `Chat.traverse_provenance` once and collect the flat visit stream or edges.
437
+ Do not print during recursion. Do not infer edge types from Ruby classes after
438
+ the fact.
439
+
440
+ ### Analysis
441
+
442
+ Load each discovered chat once. Use Chat operations for direct token entries,
443
+ totals, and any future tool-call summaries. Job token totals can be defined as
444
+ the totals of chat logs directly owned by that job; subtree totals should be
445
+ named explicitly because they include dependencies and can overlap in a DAG.
446
+
447
+ The current command sometimes presents a job total that recursively includes
448
+ its logs and descendants without making scope obvious. Reports should label
449
+ `direct` versus `subtree` totals.
450
+
451
+ ### Rendering
452
+
453
+ Tree, compact flow, DOT, and plots should consume the same node/edge arrays.
454
+ Tree indentation is a spanning-tree presentation of a DAG; repeated nodes
455
+ should be shown as references rather than recursively expanded. Flow arrow
456
+ direction should be selected by relation in rendering only.
457
+
458
+ This removes the global `@flow_*` caches and most type/path guesses from the
459
+ command.
460
+
461
+ ## How ChatAnalyst should change
462
+
463
+ ChatAnalyst does not need Session.
464
+
465
+ Each task can use one Workflow helper that returns the traversal stream or a
466
+ plain collected Hash for the duration of that task. A helper is appropriate
467
+ because it is local executable reuse, not a new domain abstraction. For
468
+ example:
469
+
470
+ helper :provenance_records do |file, warnings = []|
471
+ Chat.traverse_provenance(file, on_error: ->(...) { warnings << ... }).to_a
472
+ end
473
+
474
+ Tasks can then remain focused:
475
+
476
+ - `message_index`: iterate discovered chat files and emit address plus lineage;
477
+ - `message_content`: retrieve exact addresses;
478
+ - `chat_overview`: collect structural nodes and edges;
479
+ - `chat_tool_calls`: call the shared Chat pairing primitive;
480
+ - `chat_tokens`: call source-aware trace/token operations;
481
+ - `chat_agents`: filter paired tool calls, while marking inferred log matches as
482
+ inferred;
483
+ - `chat_report`: use Workflow dependencies on the smaller tasks if caching is
484
+ desirable, or combine their plain helper results.
485
+
486
+ `chat_report` should not rerun a hidden second traversal for each subreport
487
+ inside one task. Either collect once locally or make reports proper Workflow
488
+ dependencies. That decision is normal Workflow composition, not provenance
489
+ architecture.
490
+
491
+ ## Identity and deduplication rules
492
+
493
+ ### Chats
494
+
495
+ For one filesystem, use an expanded real path where possible. Preserve the
496
+ original reference as an alias for reporting. Do not silently merge different
497
+ files merely because their contents overlap.
498
+
499
+ ### Jobs
500
+
501
+ Use the loaded Step path as the physical identity. A logical identity such as
502
+ workflow, task, and result basename can identify mirrored `~/.scout` and
503
+ `~/.rbbt` candidates, but merging them is safe only if their relevant info and
504
+ result agree. If two physical jobs share a logical identity but disagree,
505
+ retain both and issue a warning rather than selecting one silently.
506
+
507
+ ### Messages
508
+
509
+ Use `[chat_path, index]` as the address and the existing digest as
510
+ `lineage_id`.
511
+
512
+ ### Inference events
513
+
514
+ Introduce a random locally generated `inference_id` (or equivalently named
515
+ field) in each direct meta record. Generate it once per actual backend request
516
+ and preserve it through cached replay and copied chat history. Provider request
517
+ IDs can be stored separately when available, but should not be required.
518
+ Legacy records fall back to lineage identity with an explicit
519
+ `deduplication: legacy_lineage` qualification in precise reports.
520
+
521
+ ### Tool calls
522
+
523
+ Use provider/tool call ID scoped by chat lineage or chat address. Do not assume
524
+ a provider call ID is globally unique across all files.
525
+
526
+ ## Correctness tests to add before migration
527
+
528
+ Traversal tests should use temporary persisted jobs/chats and cover:
529
+
530
+ 1. a root chat with no references;
531
+ 2. a chat importing another chat;
532
+ 3. a chat projected from a job;
533
+ 4. a job with multiple recursive log chats;
534
+ 5. a job with a chat dependency and a non-chat dependency;
535
+ 6. shared dependencies visited once but represented by all relevant edges;
536
+ 7. a cycle caused by a chat result projecting its own producer job;
537
+ 8. missing import, missing job, malformed chat, and unreadable Step info under
538
+ both strict and warning policies;
539
+ 9. two physical job roots with the same logical digest;
540
+ 10. an aborted or error Step with partial log files;
541
+ 11. traversal starting from a Step rather than a chat;
542
+ 12. deterministic traversal order.
543
+
544
+ Tracing/accounting tests should cover:
545
+
546
+ 1. copied direct inference history deduplicated once;
547
+ 2. two identical real requests with distinct `inference_id` counted twice;
548
+ 3. legacy identical records reported with fallback deduplication;
549
+ 4. job projection meta excluded from direct token totals;
550
+ 5. cache, cache-write, and reasoning token fields;
551
+ 6. source addresses retained for every segment and covered message;
552
+ 7. orphan and consecutive meta records;
553
+ 8. the same tool call ID appearing in two chat files;
554
+ 9. missing tool output versus an output containing an exception;
555
+ 10. socialized agent calls where a log association is only inferred.
556
+
557
+ Consumer tests should run `prov` and ChatAnalyst against the same fixture and
558
+ assert that they discover the same chat paths, jobs, and structural edges.
559
+
560
+ ## Concrete defects worth fixing opportunistically
561
+
562
+ These are small enough to address while introducing the primitive:
563
+
564
+ - make `job_agent_chats` actually load and return Chat values, or rename it;
565
+ - document and test whether each `job_*` helper is direct or recursive;
566
+ - remove broad rescue-to-empty behavior from structural readers;
567
+ - remove the unused `load_chat` in `prov`;
568
+ - stop detecting job roots only through `<filename>.files`;
569
+ - avoid `merge!` on a recursive result that may be `nil` for a seen node;
570
+ - require `set` explicitly in `prov` if it continues to use Set directly;
571
+ - stop classifying an agent as a distinct persistent node based solely on a
572
+ path substring;
573
+ - make token scope explicit (`direct file`, `direct job logs`, or `subtree`);
574
+ - correct maintained documentation that currently claims recursive behavior or
575
+ APIs that do not match the code.
576
+
577
+ ## Relationship to the deferred concerns
578
+
579
+ ### Exceptions, aborted managers, and active-agent monitoring
580
+
581
+ The proposed traversal already helps historical failures because it must visit
582
+ jobs regardless of `done?`, expose Step status, and discover partial log files.
583
+ It should not require a new monitoring abstraction.
584
+
585
+ If active agents later write state under `var`, the clean integration is to
586
+ make that state another persisted artifact referenced by a Step or chat, not to
587
+ put live-agent state into the provenance walker. Historical provenance and
588
+ live monitoring have different consistency requirements.
589
+
590
+ Backend failure snapshots currently written to anonymous TmpFile paths are
591
+ weak provenance because the path is logged but not structurally linked to the
592
+ Step/chat. A future change could save the snapshot under the owning job's
593
+ files directory when `agent.job` is available, or record its path in Step info.
594
+ That would make it discoverable without changing traversal semantics. This is
595
+ not required for the provenance refactor.
596
+
597
+ ### `ask_user`
598
+
599
+ An `ask_user` operation should appear as an ordinary tool call. Its spool item
600
+ and eventual response can carry the tool call ID or a derived request ID. The
601
+ generic tool-call pairing and provenance addresses proposed here are enough;
602
+ the walker should not special-case human interaction.
603
+
604
+ ### Agents declaring impossible work
605
+
606
+ An explicit failure should become either a tool output containing structured
607
+ failure or a normal Workflow/Step exception and status. Again, traversal only
608
+ needs to preserve and expose that evidence. It should not decide whether the
609
+ failure was justified.
610
+
611
+ ## Recommended implementation sequence
612
+
613
+ 1. **Specify and test direct relations.** Correct helper names/return types and
614
+ add fixtures for chats, jobs, dependencies, logs, and imports.
615
+ 2. **Add `Chat.traverse_provenance`.** Keep it procedural, block-oriented, and
616
+ explicit about errors.
617
+ 3. **Reimplement existing collectors from traversal.** Preserve compatibility
618
+ where inexpensive.
619
+ 4. **Migrate `prov`.** Separate discovery, analysis, and rendering; compare
620
+ output against current real jobs.
621
+ 5. **Add source-aware tracing and tool-call pairing.** Keep plain Hash/Array
622
+ results.
623
+ 6. **Migrate ChatAnalyst and delete Session.** Use Workflow helpers and tasks,
624
+ not replacement classes.
625
+ 7. **Add persisted inference IDs.** Update accounting to prefer them while
626
+ retaining legacy lineage fallback.
627
+ 8. **Revise maintained developer documentation.** Keep this investigation as
628
+ the detailed rationale and document only the stable concepts/API in
629
+ `doc/developer/Provenance.md`.
630
+
631
+ ## Final design rule
632
+
633
+ A useful test for every proposed provenance abstraction is:
634
+
635
+ > Does this represent a persisted fact in Chat or Workflow, or is it only state
636
+ > needed while producing one report?
637
+
638
+ Persisted facts belong in Chat/Step primitives. Temporary report state belongs
639
+ in local Arrays, Hashes, Sets, and blocks. If an object exists only to hold a
640
+ queue, a seen set, loaded chats, edges, and warnings, it should not be a class.