scout-ai 1.2.3 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (174) hide show
  1. checksums.yaml +4 -4
  2. data/.vimproject +138 -50
  3. data/README.md +171 -290
  4. data/Rakefile +17 -1
  5. data/VERSION +1 -1
  6. data/doc/Improvements.md +325 -0
  7. data/doc/StartHere.md +110 -0
  8. data/doc/developer/Architecture.md +126 -0
  9. data/doc/developer/Backends.md +199 -0
  10. data/doc/developer/ChatLifecycle.md +183 -0
  11. data/doc/developer/DelegationInternals.md +295 -0
  12. data/doc/developer/DesignPrinciples.md +245 -0
  13. data/doc/developer/PromptProcessing.md +292 -0
  14. data/doc/developer/Provenance.md +317 -0
  15. data/doc/user/BuildingAgents.md +345 -0
  16. data/doc/user/Cookbook.md +333 -0
  17. data/doc/user/CoreConcepts.md +181 -0
  18. data/doc/user/Delegation.md +191 -0
  19. data/doc/user/GettingStarted.md +159 -0
  20. data/doc/user/ManagingContext.md +163 -0
  21. data/doc/user/MultiAgentWorkflows.md +256 -0
  22. data/doc/user/Python.md +159 -0
  23. data/doc/user/RunningInference.md +200 -0
  24. data/doc/user/ToolCalling.md +193 -0
  25. data/doc/user/WritingChats.md +197 -0
  26. data/lib/scout/llm/agent/chat.rb +61 -11
  27. data/lib/scout/llm/agent/delegate.rb +274 -65
  28. data/lib/scout/llm/agent/iterate.rb +2 -2
  29. data/lib/scout/llm/agent/save.rb +273 -0
  30. data/lib/scout/llm/agent/workflow.rb +164 -0
  31. data/lib/scout/llm/agent.rb +86 -61
  32. data/lib/scout/llm/ask.rb +62 -17
  33. data/lib/scout/llm/backends/anthropic.rb +9 -2
  34. data/lib/scout/llm/backends/bedrock.rb +15 -3
  35. data/lib/scout/llm/backends/default.rb +183 -99
  36. data/lib/scout/llm/backends/glm.rb +58 -0
  37. data/lib/scout/llm/backends/huggingface.rb +196 -26
  38. data/lib/scout/llm/backends/ollama.rb +13 -1
  39. data/lib/scout/llm/backends/openai.rb +0 -2
  40. data/lib/scout/llm/backends/openwebui.rb +20 -13
  41. data/lib/scout/llm/backends/relay.rb +22 -22
  42. data/lib/scout/llm/backends/responses.rb +1 -1
  43. data/lib/scout/llm/chat/agent_meta.rb +264 -0
  44. data/lib/scout/llm/chat/annotation.rb +39 -10
  45. data/lib/scout/llm/chat/parse.rb +28 -6
  46. data/lib/scout/llm/chat/persist.rb +25 -0
  47. data/lib/scout/llm/chat/process/clear.rb +41 -6
  48. data/lib/scout/llm/chat/process/files.rb +21 -6
  49. data/lib/scout/llm/chat/process/meta.rb +421 -34
  50. data/lib/scout/llm/chat/process/options.rb +21 -1
  51. data/lib/scout/llm/chat/process/tools.rb +56 -15
  52. data/lib/scout/llm/chat/process.rb +4 -0
  53. data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
  54. data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
  55. data/lib/scout/llm/chat/prompt.rb +48 -0
  56. data/lib/scout/llm/chat/provenance.rb +775 -0
  57. data/lib/scout/llm/chat/tool_calls.rb +76 -0
  58. data/lib/scout/llm/chat.rb +18 -2
  59. data/lib/scout/llm/embed.rb +11 -3
  60. data/lib/scout/llm/image.rb +86 -0
  61. data/lib/scout/llm/mcp.rb +10 -2
  62. data/lib/scout/llm/rag.rb +3 -3
  63. data/lib/scout/llm/tools/call.rb +160 -11
  64. data/lib/scout/llm/tools/knowledge_base.rb +1 -1
  65. data/lib/scout/llm/tools/workflow.rb +32 -16
  66. data/lib/scout/model/python/huggingface/causal.rb +23 -5
  67. data/lib/scout/model/python/huggingface.rb +2 -1
  68. data/lib/scout-ai.rb +1 -0
  69. data/python/README.md +197 -14
  70. data/python/scout_ai/huggingface/eval.py +245 -34
  71. data/python/tests/test_huggingface_eval.py +58 -0
  72. data/research/ChatAnalyst-required-changes.md +167 -0
  73. data/research/agent-delegation-analysis.md +810 -0
  74. data/research/agent-meta-provenance-integration-plan.md +622 -0
  75. data/research/agent-workflow-analysis.md +1120 -0
  76. data/research/backends-analysis.md +836 -0
  77. data/research/chat-core-analysis.md +946 -0
  78. data/research/chatanalyst-provenance/00-baseline.md +30 -0
  79. data/research/chatanalyst-provenance/01-repo-map.md +60 -0
  80. data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
  81. data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
  82. data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
  83. data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
  84. data/research/chatanalyst-provenance/07-critic-review.md +25 -0
  85. data/research/chatanalyst-provenance/final-report.md +45 -0
  86. data/research/chatanalyst-provenance/resumption.md +37 -0
  87. data/research/coding-philosophy-analysis.md +928 -0
  88. data/research/commands-analysis.md +947 -0
  89. data/research/multi-agent-patterns-analysis.md +853 -0
  90. data/research/prompt-strategies-analysis.md +630 -0
  91. data/research/prov-verbosity-fix-notes.md +77 -0
  92. data/research/provenance-analysis.md +469 -0
  93. data/research/provenance-navigation-design.md +640 -0
  94. data/research/synthesis-report.md +487 -0
  95. data/research/tools-system-analysis.md +779 -0
  96. data/scout-ai.gemspec +100 -11
  97. data/scout_commands/agent/ask +13 -3
  98. data/scout_commands/agent/kb +2 -0
  99. data/scout_commands/llm/ask +11 -4
  100. data/scout_commands/llm/md +76 -0
  101. data/scout_commands/llm/process_queries +48 -0
  102. data/scout_commands/llm/prov +602 -0
  103. data/scout_commands/llm/word +71 -0
  104. data/scout_commands/workflow/mcp +43 -0
  105. data/share/word/reference.docx +0 -0
  106. data/test/etc/AI/mock.yaml +11 -0
  107. data/test/fixtures/backends/anthropic.json +19 -0
  108. data/test/fixtures/backends/anthropic_tool_use.json +24 -0
  109. data/test/fixtures/backends/bedrock.json +8 -0
  110. data/test/fixtures/backends/bedrock_embedding.json +3 -0
  111. data/test/fixtures/backends/bedrock_tool_use.json +17 -0
  112. data/test/fixtures/backends/ollama.json +16 -0
  113. data/test/fixtures/backends/ollama_tool_call.json +27 -0
  114. data/test/fixtures/backends/openai_chat.json +21 -0
  115. data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
  116. data/test/fixtures/backends/responses.json +33 -0
  117. data/test/fixtures/backends/responses_tool_call.json +28 -0
  118. data/test/integration/README.md +32 -0
  119. data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
  120. data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
  121. data/test/integration/scout/llm/backends/test_relay.rb +52 -0
  122. data/test/integration/scout/llm/test_infrastructure.rb +74 -0
  123. data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
  124. data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
  125. data/test/integration/scout/model/test_base.rb +91 -0
  126. data/test/scout/llm/agent/test_chat.rb +8 -2
  127. data/test/scout/llm/agent/test_save.rb +413 -0
  128. data/test/scout/llm/agent/test_workflow.rb +110 -0
  129. data/test/scout/llm/backends/test_anthropic.rb +93 -10
  130. data/test/scout/llm/backends/test_bedrock.rb +118 -2
  131. data/test/scout/llm/backends/test_huggingface.rb +137 -42
  132. data/test/scout/llm/backends/test_ollama.rb +70 -20
  133. data/test/scout/llm/backends/test_openwebui.rb +42 -40
  134. data/test/scout/llm/backends/test_relay.rb +4 -2
  135. data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
  136. data/test/scout/llm/chat/process/test_meta.rb +518 -0
  137. data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
  138. data/test/scout/llm/chat/test_agent_meta.rb +357 -0
  139. data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
  140. data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
  141. data/test/scout/llm/chat/test_parse.rb +70 -15
  142. data/test/scout/llm/chat/test_prov_cli.rb +274 -0
  143. data/test/scout/llm/chat/test_provenance.rb +240 -0
  144. data/test/scout/llm/chat/test_tool_calls.rb +38 -0
  145. data/test/scout/llm/test_agent.rb +13 -36
  146. data/test/scout/llm/test_ask.rb +75 -52
  147. data/test/scout/llm/test_chat.rb +107 -13
  148. data/test/scout/llm/test_embed.rb +48 -0
  149. data/test/scout/llm/test_rag.rb +23 -16
  150. data/test/scout/llm/test_tools.rb +12 -1
  151. data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
  152. data/test/scout/llm/tools/test_mcp.rb +5 -3
  153. data/test/scout/llm/tools/test_workflow.rb +23 -2
  154. data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
  155. data/test/scout/model/python/huggingface/test_causal.rb +9 -3
  156. data/test/scout/model/python/huggingface/test_classification.rb +11 -2
  157. data/test/scout/model/python/test_torch.rb +2 -0
  158. data/test/scout/model/python/torch/test_helpers.rb +4 -0
  159. data/test/scout/model/test_base.rb +4 -2
  160. data/test/support/availability.rb +231 -0
  161. data/test/support/fake_clients.rb +138 -0
  162. data/test/support/fixtures.rb +21 -0
  163. data/test/support/infrastructure_probes.rb +136 -0
  164. data/test/support/mock_backend.rb +215 -0
  165. data/test/test_helper.rb +32 -2
  166. metadata +99 -10
  167. data/doc/Agent.md +0 -327
  168. data/doc/Chat.md +0 -458
  169. data/doc/LLM.md +0 -340
  170. data/doc/RAG.md +0 -129
  171. data/scout_commands/documenter +0 -148
  172. data/test/scout/llm/backends/test_openai.rb +0 -192
  173. data/test/scout/llm/backends/test_responses.rb +0 -238
  174. data/test/scout/llm/test_parse.rb +0 -98
@@ -0,0 +1,622 @@
1
+ # Agent-meta provenance integration: implementation and test plan
2
+
3
+ > **Superseded (historical document).** Commit `efd8ebbc` ("Change how meta is
4
+ > passed in function_call_output. Auto-save agent chats in .files") replaced
5
+ > the serialized `agent_meta` writer shape described here with a `meta` key
6
+ > holding an Array of already-deserialized field Hashes. The reader still
7
+ > accepts both formats; see `doc/developer/Provenance.md` for current
8
+ > behavior. The rest of this file is kept verbatim as the design record.
9
+
10
+ > This is an implementation report for Scout-AI and ChatAnalyst coding agents.
11
+ > It describes a change to provenance handling for `agent_meta` receipts that
12
+ > are embedded in `function_call_output` records. It does not introduce a
13
+ > Session, ChatGraph, ProvenanceContext, or database abstraction.
14
+
15
+ ## Decision summary
16
+
17
+ Scout-AI should treat `agent_meta` as **embedded provenance evidence**.
18
+
19
+ It must be included in two places:
20
+
21
+ 1. **Structural traversal** must follow `job=` values found in an `agent_meta`
22
+ receipt, so a delegated agent whose answer is produced by a `chat_task`
23
+ exposes the actual producer Step, dependencies, logs, and result.
24
+ 2. **Token accounting** must include direct token meta records found in
25
+ `agent_meta`, so socialized delegation remains auditable even when there is
26
+ no separately saved child chat.
27
+
28
+ It must not be represented as a third graph-node kind, converted into ordinary
29
+ parent-chat `meta:` messages, or assigned its own stateful context object.
30
+
31
+ The saved child chat and the receipt may both contain the same direct inference
32
+ metadata. They are two evidence locations for one inference event. Deduplicate
33
+ by `inference_id`, preserve both locations for auditing, and count the event
34
+ once.
35
+
36
+ ## Current behavior and gap
37
+
38
+ `LLM.process_calls` already creates the required receipt. When a tool returns
39
+ an `LLM::Agent`, it:
40
+
41
+ 1. runs `agent.chat(return_messages: true)`;
42
+ 2. appends the resulting messages to the child agent's current chat;
43
+ 3. extracts `Chat.find_role(res, :meta)`;
44
+ 4. serializes those records as `agent_meta` in the parent
45
+ `function_call_output` JSON.
46
+
47
+ The receipt shape is therefore intentionally narrow and stable:
48
+
49
+ function_call_output: {
50
+ "name": "ask",
51
+ "content": "child answer",
52
+ "id": "call_...",
53
+ "agent_meta": [
54
+ {"role": "meta", "content": "pt=... tt=... inference_id=..."},
55
+ {"role": "meta", "content": "job=Workflow/ask/..."}
56
+ ],
57
+ "start_timestamp": "...",
58
+ "timestamp": "..."
59
+ }
60
+
61
+ The current provenance walker only reads ordinary persisted Chat `meta:`
62
+ messages through `chat.jobs`. It does not parse `function_call_output` JSON,
63
+ so it misses `job=` records stored in `agent_meta`.
64
+
65
+ Likewise, `Chat.token_totals` and the current `llm prov` aggregate direct
66
+ messages found in discovered chat files. They do not see direct token events
67
+ inside a receipt. ChatAnalyst can recover this only through specialized agent
68
+ instructions, which is exactly the framework responsibility that should move
69
+ into Scout-AI primitives.
70
+
71
+ ## Semantics and invariants
72
+
73
+ ### The three relevant persisted facts
74
+
75
+ | Fact | Authoritative use |
76
+ |---|---|
77
+ | Parent function call/output | The parent asked a tool or agent and received this result. |
78
+ | `agent_meta` receipt | Portable evidence of direct child inference events and child producer jobs. It is especially important when no child log is discoverable. |
79
+ | Child chat / Step logs | Full child execution history: messages, tool calls, direct inference traces, dependencies, files, statuses, and exceptions. |
80
+
81
+ The receipt is not a replacement for a saved child log. It is a compact
82
+ projection of some of the same evidence.
83
+
84
+ ### What is authoritative for token counting
85
+
86
+ The authoritative counted unit is a **direct inference event**, identified by
87
+ `inference_id`.
88
+
89
+ - A direct meta record in a saved child `agent.chat` and the same record in an
90
+ `agent_meta` receipt represent one event if their `inference_id` matches.
91
+ - Both source locations must remain visible in an audit report.
92
+ - The event's token fields are counted once.
93
+ - A `job=` meta record is not a token event. It is a producer reference and
94
+ must cause traversal into the referenced Step.
95
+ - `*_c` and `*_s` remain checkpoints and must never be summed.
96
+
97
+ ### Why `inference_id` is specifically justified here
98
+
99
+ The receipt/log duplication is a real and common ambiguity, unlike a merely
100
+ hypothetical retry:
101
+
102
+ - parent output contains the Worker’s two direct metas;
103
+ - Worker `agent.chat`, when saved, contains the same two metas;
104
+ - both are valid persisted evidence;
105
+ - only the event identity says they are the same paid requests.
106
+
107
+ Without `inference_id`, no exact general rule can distinguish a receipt copy
108
+ from an independent request with matching metadata. The field is therefore a
109
+ join key between two persisted representations of one actual inference.
110
+
111
+ ### Never inject receipt records into the parent Chat
112
+
113
+ Do not append `agent_meta` records to the parent Chat as ordinary `role: meta`
114
+ messages.
115
+
116
+ That would corrupt parent conversational bookkeeping:
117
+
118
+ - `Chat.meta` selects checkpoints used by later parent requests;
119
+ - child `*_c` values belong to the child conversation;
120
+ - parent and child session totals must remain separate;
121
+ - a delegated receipt is evidence attached to a tool output, not a new parent
122
+ response segment.
123
+
124
+ Extraction must be read-only and happen only in provenance analysis.
125
+
126
+ ## Proposed Scout-AI primitives
127
+
128
+ The implementation should be small, procedural, and return ordinary Arrays
129
+ and Hashes.
130
+
131
+ ### 1. Parse agent-meta evidence from tool outputs
132
+
133
+ Add a Chat helper near `lib/scout/llm/chat/tool_calls.rb`, reusing
134
+ `Chat.tool_calls` to avoid reparsing output pairing logic.
135
+
136
+ Suggested API:
137
+
138
+ Chat.agent_meta_evidence(chat, source: nil)
139
+
140
+ It returns one plain Hash per valid receipt meta record, for example:
141
+
142
+ {
143
+ origin: :agent_meta,
144
+ meta: <IndiferentHash parsed by Chat.parse_meta>,
145
+ source: "/path/to/parent.chat",
146
+ output_address: ["/path/to/parent.chat", 17],
147
+ evidence_address: ["/path/to/parent.chat", 17, :agent_meta, 0],
148
+ call_id: "call_123",
149
+ tool_name: "ask",
150
+ agent_meta_index: 0,
151
+ raw_message: {role: "meta", content: "..."}
152
+ }
153
+
154
+ `source` is optional, matching existing source-aware Chat APIs. If omitted,
155
+ the address can use only message indexes.
156
+
157
+ Extraction rules:
158
+
159
+ - inspect parsed `function_call_output` records, not raw text with regular
160
+ expressions;
161
+ - accept an `agent_meta` Array only;
162
+ - accept entries only when they are Hashes with `role == "meta"` and a String
163
+ `content`;
164
+ - parse content through `Chat.parse_meta`;
165
+ - retain malformed receipt entries as warnings when the caller asks for
166
+ warnings, but never silently reinterpret arbitrary output JSON as
167
+ provenance;
168
+ - do not limit the logic to tool name `ask`: the envelope is safe to support
169
+ generically, although reports may label agent-oriented tools specially.
170
+
171
+ ### 2. Expose all meta evidence without altering ordinary Chat semantics
172
+
173
+ Add:
174
+
175
+ Chat.meta_evidence(chat, source: nil)
176
+
177
+ This returns normal persisted `role: meta` records plus
178
+ `Chat.agent_meta_evidence` records, with a shared shape that includes
179
+ `origin`.
180
+
181
+ Suggested origins:
182
+
183
+ - `:chat_meta` for a normal Chat message;
184
+ - `:agent_meta` for a receipt embedded in a function output.
185
+
186
+ Do not change `Chat#meta`, `chat.role_messages(:meta)`, or the current
187
+ meaning of `Chat.token_totals([chat])`. Those APIs describe the local Chat and
188
+ must not unexpectedly absorb delegated receipts.
189
+
190
+ ### 3. Expose embedded producer-job references
191
+
192
+ Add:
193
+
194
+ Chat.agent_meta_job_references(chat, source: nil)
195
+
196
+ This filters `agent_meta_evidence` to records with `meta[:job]`. It returns the
197
+ reference and the evidence location.
198
+
199
+ The existing `chat.jobs` can remain a local-chat API. Do not silently redefine
200
+ it to include nested receipt jobs; callers need to distinguish direct visible
201
+ projection from delegated receipt provenance.
202
+
203
+ ### 4. Add a provenance-aware token-event collector
204
+
205
+ Add a root-oriented API, for example:
206
+
207
+ Chat.provenance_token_events(root, **traversal_options)
208
+ Chat.provenance_token_totals(root, **traversal_options)
209
+
210
+ This API should:
211
+
212
+ 1. invoke `Chat.traverse_provenance` to discover persisted chats and jobs;
213
+ 2. load every discovered chat once;
214
+ 3. collect ordinary direct events from source-aware chat tracing;
215
+ 4. collect direct receipt events from `agent_meta_evidence`;
216
+ 5. discard projection records with `meta[:job]` from token addition;
217
+ 6. group direct events by `inference_id`;
218
+ 7. return one event record with all evidence locations;
219
+ 8. sum canonical direct fields only: `pt`, `ct`, `tt`, `cct`, `cwt`, and `rt`.
220
+
221
+ A returned event should include enough audit information to explain a total:
222
+
223
+ {
224
+ inference_id: "...",
225
+ meta: <canonical parsed direct meta>,
226
+ tokens: {pt: ..., ct: ..., tt: ...},
227
+ evidence: [
228
+ {origin: :agent_meta, evidence_address: [...]},
229
+ {origin: :chat_meta, meta_address: [...]}
230
+ ],
231
+ deduplication: :inference_id
232
+ }
233
+
234
+ ### Legacy records without `inference_id`
235
+
236
+ Do not claim exact receipt/log deduplication for old data.
237
+
238
+ For ordinary persisted chat meta, preserve the existing lineage-based fallback
239
+ already used by `Chat.trace_chats`.
240
+
241
+ For an embedded receipt meta without an inference ID, there is no full child
242
+ conversation in which to compute its lineage. Use a receipt-address-based
243
+ fallback and mark it clearly, for example:
244
+
245
+ deduplication: :receipt_unresolved
246
+
247
+ If a provider response ID is present and Scout-AI considers it stable enough,
248
+ it may be used as an explicit secondary identity. Do not silently use
249
+ heuristic equality of token counts, timestamps, or reasoning text as proof of
250
+ identity.
251
+
252
+ Reports should warn that legacy receipt/log records can be overcounted when no
253
+ shared inference or provider identity exists.
254
+
255
+ ### Identity conflicts
256
+
257
+ If two records share an `inference_id` but disagree on direct token fields,
258
+ provider response ID, or other immutable event facts:
259
+
260
+ - do not sum both;
261
+ - retain both evidence records;
262
+ - emit a provenance conflict warning;
263
+ - make the conflict visible in `prov` and ChatAnalyst.
264
+
265
+ This detects a broken producer rather than hiding it behind deduplication.
266
+
267
+ ## Traversal changes
268
+
269
+ ### Add `:agent_job` as a relation
270
+
271
+ Extend `Chat::PROVENANCE_RELATIONS` with `:agent_job`.
272
+
273
+ When visiting a Chat node, the walker should enqueue both:
274
+
275
+ - ordinary direct `chat.jobs` as `:job` edges;
276
+ - receipt `agent_meta_job_references` as `:agent_job` edges.
277
+
278
+ The parent remains the enclosing Chat node and the child remains a native
279
+ Step. There is no receipt node.
280
+
281
+ This relation is useful because it preserves why the job was discovered:
282
+
283
+ | Relation | Meaning |
284
+ |---|---|
285
+ | `job` | A visible response segment in this chat was projected from the job. |
286
+ | `agent_job` | A delegated agent receipt inside a tool output says its response was projected from the job. |
287
+
288
+ The referenced Step then follows normal `dependency`, `log`, and `result`
289
+ relations. Existing cycle and shared-node handling applies unchanged.
290
+
291
+ ### Error handling
292
+
293
+ Malformed receipt JSON or unreadable receipt job references must be reported
294
+ through the existing `on_error`/warning path with:
295
+
296
+ - enclosing chat path;
297
+ - tool output address;
298
+ - call ID and tool name when available;
299
+ - relation `:agent_job`;
300
+ - original reference.
301
+
302
+ A malformed receipt should not hide other normal provenance from the same
303
+ chat.
304
+
305
+ ### Do not follow ordinary imports
306
+
307
+ The current traversal intentionally treats imports as compilation composition,
308
+ not persisted provenance edges. Preserve that policy. The agent-meta work does
309
+ not require reintroducing import traversal.
310
+
311
+ ## `scout-ai llm prov` changes
312
+
313
+ The command already separates graph discovery from token computation. Update
314
+ only those layers.
315
+
316
+ ### Graph discovery and rendering
317
+
318
+ 1. Call the enhanced traversal.
319
+ 2. Include `:agent_job` in adjacency sorting, after direct `:job` and before
320
+ ordinary dependency/log display as appropriate.
321
+ 3. In default tree mode, show a job discovered through a receipt with an
322
+ explicit label, for example:
323
+
324
+ chat Manager.chat
325
+ delegated-job Worker/ask abc12345
326
+
327
+ The job remains a normal job node; `delegated-job` is the edge label.
328
+
329
+ 4. In flow and DOT modes, reverse `:agent_job` in the same way as `:job`,
330
+ because the job produces content used by the parent chat. Use a distinct
331
+ visual style or label such as `delegated_result` so it is not confused with
332
+ a direct projected job result.
333
+
334
+ 5. Do not make a Graphviz node for every receipt or every direct inference.
335
+ The graph should remain a Chat/Step structural graph.
336
+
337
+ ### Token display
338
+
339
+ Replace direct uses of `Chat.token_totals(chats)` in `prov` aggregate logic
340
+ with `Chat.provenance_token_totals` or the equivalent event collector.
341
+
342
+ The default total must include receipt-only child usage. It must not add it
343
+ again when the same child log is reachable.
344
+
345
+ The current `--component` mode should become explicit about source scope. The
346
+ recommended output distinctions are:
347
+
348
+ - `local`: direct meta events physically stored in the chat/log;
349
+ - `receipt`: child direct events found only or also in `agent_meta` receipts;
350
+ - `aggregate`: deduplicated union of all reachable event identities.
351
+
352
+ For a compact tree, do not print a separate line for every receipt by default.
353
+ Instead, append a concise annotation to the parent chat or tool call summary,
354
+ for example:
355
+
356
+ delegated receipt: 2 events, total=17.0k, Worker/test_sum
357
+
358
+ When `--component` is enabled, print receipt components individually with call
359
+ ID and source address.
360
+
361
+ ### New CLI option
362
+
363
+ Add a focused diagnostic option rather than overloading flow output. Suggested
364
+ name:
365
+
366
+ --evidence
367
+
368
+ It prints the deduplicated direct inference event table:
369
+
370
+ | Inference ID | Tokens | Evidence | Status |
371
+ |---|---:|---|---|
372
+ | `eed7eb...` | 8463 | parent ask output; Worker agent.chat | counted once |
373
+ | `a477d8...` | 8550 | parent ask output; Worker agent.chat | counted once |
374
+
375
+ It should also display:
376
+
377
+ - receipt-only events;
378
+ - legacy unresolved events;
379
+ - inference-ID conflicts;
380
+ - job projection references, but with no direct tokens.
381
+
382
+ This is the right command for diagnosing why a total contains child work. The
383
+ normal tree and flow should remain compact.
384
+
385
+ ## ChatAnalyst changes
386
+
387
+ ChatAnalyst should consume the Scout-AI primitives and delete any bespoke
388
+ receipt-parsing logic it currently has. It should not need special agent
389
+ instructions to discover `agent_meta`.
390
+
391
+ ### `chat_overview`
392
+
393
+ Add:
394
+
395
+ - `agent_job` structural edges;
396
+ - number of receipt records per chat;
397
+ - number of receipt-only direct events;
398
+ - provenance warnings for malformed receipts or conflicts.
399
+
400
+ ### `chat_tool_calls`
401
+
402
+ For each function call output, add a compact receipt summary:
403
+
404
+ agent_meta: {
405
+ direct_events: 2,
406
+ job_references: ["Worker/ask/..."],
407
+ token_total: {tt: 17013},
408
+ event_ids: ["eed7...", "a477..."]
409
+ }
410
+
411
+ Keep the raw answer content and tool success status separate from token
412
+ attribution.
413
+
414
+ ### `chat_tokens`
415
+
416
+ This task should switch from plain `Chat.token_totals` to the new provenance
417
+ event collector. Return:
418
+
419
+ - aggregate deduplicated totals;
420
+ - `events` or a compact per-event index;
421
+ - per-chat local totals;
422
+ - receipt-derived totals;
423
+ - receipt-only totals;
424
+ - duplicate evidence count;
425
+ - unresolved legacy receipt count;
426
+ - identity conflict warnings.
427
+
428
+ Do not describe a receipt contribution as a second paid inference when its ID
429
+ also appears in a saved child log.
430
+
431
+ ### `chat_agents`
432
+
433
+ An agent interaction should report:
434
+
435
+ - call ID;
436
+ - target agent and conversation when present in arguments;
437
+ - receipt event IDs and totals;
438
+ - linked `agent_job` Steps, if any;
439
+ - whether child evidence was receipt-only, log-only, or both;
440
+ - whether the log association is structural, inferred by naming convention, or
441
+ absent.
442
+
443
+ ### `message_index` and `message_content`
444
+
445
+ Do not pretend embedded receipt metas are ordinary child-chat messages.
446
+
447
+ Either:
448
+
449
+ 1. add a dedicated `meta_evidence` task; or
450
+ 2. allow `message_index --role meta` to include an `origin` field and a nested
451
+ receipt address.
452
+
453
+ A dedicated task is clearer. It can return normal and embedded meta evidence
454
+ without asserting that an embedded record has a full child message history.
455
+
456
+ Suggested address format:
457
+
458
+ [parent_chat_path, function_output_index, :agent_meta, meta_index]
459
+
460
+ `message_content` may resolve this address by re-parsing the parent output;
461
+ there is no separate physical child message at that address.
462
+
463
+ ### `chat_report`
464
+
465
+ Include concise highlights only:
466
+
467
+ - total deduplicated direct tokens;
468
+ - local versus receipt-only contribution;
469
+ - number of event IDs with multiple evidence locations;
470
+ - count of `agent_job` edges;
471
+ - conflicts and unresolved legacy receipts.
472
+
473
+ ## Test plan
474
+
475
+ All fixtures must be offline and use persisted chat text plus temporary job
476
+ layouts. Do not call model providers.
477
+
478
+ ### A. Receipt-only socialized delegation
479
+
480
+ Fixture:
481
+
482
+ - parent chat has one `ask` function call/output;
483
+ - output includes two direct `agent_meta` records with distinct inference IDs;
484
+ - no saved Worker chat or Worker job exists.
485
+
486
+ Assertions:
487
+
488
+ - traversal still has only the parent Chat node;
489
+ - provenance token events include both Worker events;
490
+ - aggregate total includes parent plus Worker tokens;
491
+ - `prov --evidence` identifies both events as receipt-only;
492
+ - ChatAnalyst `chat_agents` reports the receipt and its total.
493
+
494
+ ### B. Receipt plus saved Worker log
495
+
496
+ Fixture extends A with a discoverable Worker job/log containing the same two
497
+ metadata records and inference IDs.
498
+
499
+ Assertions:
500
+
501
+ - traversal discovers the Worker Step and log through normal structure;
502
+ - token total is identical to fixture A plus any additional known worker log
503
+ events, not doubled;
504
+ - event records retain both receipt and log source locations;
505
+ - `prov` default total and ChatAnalyst aggregate total agree;
506
+ - component output labels receipt evidence rather than charging it twice.
507
+
508
+ ### C. Receipt with `job=` projection
509
+
510
+ Fixture:
511
+
512
+ - parent output contains `agent_meta` with `job=Worker/ask/...`;
513
+ - Worker Step has a normal agent log with direct token metadata.
514
+
515
+ Assertions:
516
+
517
+ - traversal emits an `:agent_job` edge;
518
+ - Worker dependencies and logs are recursively reached;
519
+ - the receipt job meta itself contributes zero direct tokens;
520
+ - actual Worker direct log tokens are counted once;
521
+ - tree, flow, and DOT label the delegated producer relationship.
522
+
523
+ ### D. Nested receipt chain
524
+
525
+ Fixture:
526
+
527
+ - Manager receipt points to Worker;
528
+ - Worker log contains a second socialized receipt for Critic;
529
+ - Critic has either direct receipt-only events or a `chat_task` producer job.
530
+
531
+ Assertions:
532
+
533
+ - recursion terminates safely;
534
+ - all direct event IDs appear once in aggregate accounting;
535
+ - graph edges preserve both delegation paths;
536
+ - no Session-like recursive state object is needed.
537
+
538
+ ### E. Malformed and incomplete receipts
539
+
540
+ Fixtures:
541
+
542
+ - invalid outer function-output JSON;
543
+ - `agent_meta` is not an Array;
544
+ - entry lacks role/content;
545
+ - malformed meta content;
546
+ - unresolved `job=` path.
547
+
548
+ Assertions:
549
+
550
+ - normal chat/job provenance remains available;
551
+ - warnings include source address and call ID when available;
552
+ - no malformed record is accidentally counted as tokens;
553
+ - strict and warning callback modes behave as documented.
554
+
555
+ ### F. Identity conflict
556
+
557
+ Fixture contains two evidence locations with the same `inference_id` but
558
+ different `tt` or provider response ID.
559
+
560
+ Assertions:
561
+
562
+ - total does not silently add both;
563
+ - event records retain both facts;
564
+ - warning is emitted in `prov --evidence` and ChatAnalyst;
565
+ - automated tests make the chosen conflict policy explicit.
566
+
567
+ ### G. Legacy records
568
+
569
+ Fixture has ordinary and receipt metadata without inference IDs.
570
+
571
+ Assertions:
572
+
573
+ - normal Chat records retain current lineage fallback behavior;
574
+ - receipt records are marked unresolved unless an explicit secondary identity
575
+ is available;
576
+ - reports do not claim exact receipt/log deduplication;
577
+ - no regression occurs for old chats without `agent_meta`.
578
+
579
+ ### H. Output truncation
580
+
581
+ Fixture uses an oversized agent answer whose parent output content is replaced
582
+ with the standard truncation exception but whose `agent_meta` remains present.
583
+
584
+ Assertions:
585
+
586
+ - receipt token events remain discoverable;
587
+ - truncation state is reported separately from model cost;
588
+ - no child tokens are lost merely because parent context omitted the full text.
589
+
590
+ ## Acceptance criteria
591
+
592
+ The implementation is complete when:
593
+
594
+ 1. a root chat with only an `ask` receipt reports delegated direct token usage;
595
+ 2. a receipt plus saved child log counts each `inference_id` once;
596
+ 3. receipt `job=` entries create discoverable normal Steps through
597
+ `:agent_job` traversal edges;
598
+ 4. parent Chat checkpoint semantics remain unchanged;
599
+ 5. existing `Chat.token_totals([chat])` remains local-Chat compatible;
600
+ 6. `llm prov`, ChatAnalyst, and the new core collector return the same
601
+ deduplicated aggregate total for shared fixtures;
602
+ 7. every receipt-derived amount is traceable to a parent output address and
603
+ call ID;
604
+ 8. malformed receipts and identity conflicts are warnings, not silent
605
+ undercounting or double counting;
606
+ 9. no wrapper class is introduced merely to hold traversal or accounting state;
607
+ 10. all new tests run offline.
608
+
609
+ ## Recommended delivery order
610
+
611
+ 1. Add receipt/meta evidence extraction and unit tests.
612
+ 2. Add embedded job-reference extraction and `:agent_job` traversal tests.
613
+ 3. Add root-oriented event collection and exact inference-ID deduplication.
614
+ 4. Update `Chat.tokens` and any provenance-specific aggregate callers to use
615
+ the new collector, while retaining local `Chat.token_totals` semantics.
616
+ 5. Update `scout-ai llm prov`, including `--evidence` and component labels.
617
+ 6. Update ChatAnalyst to consume the core APIs and remove bespoke receipt
618
+ recovery logic.
619
+ 7. Add cross-consumer fixtures asserting equal totals and equivalent job
620
+ discovery.
621
+ 8. Update maintained developer documentation and preserve this plan as the
622
+ deeper design record.