scout-ai 1.2.3 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (174) hide show
  1. checksums.yaml +4 -4
  2. data/.vimproject +138 -50
  3. data/README.md +171 -290
  4. data/Rakefile +17 -1
  5. data/VERSION +1 -1
  6. data/doc/Improvements.md +325 -0
  7. data/doc/StartHere.md +110 -0
  8. data/doc/developer/Architecture.md +126 -0
  9. data/doc/developer/Backends.md +199 -0
  10. data/doc/developer/ChatLifecycle.md +183 -0
  11. data/doc/developer/DelegationInternals.md +295 -0
  12. data/doc/developer/DesignPrinciples.md +245 -0
  13. data/doc/developer/PromptProcessing.md +292 -0
  14. data/doc/developer/Provenance.md +317 -0
  15. data/doc/user/BuildingAgents.md +345 -0
  16. data/doc/user/Cookbook.md +333 -0
  17. data/doc/user/CoreConcepts.md +181 -0
  18. data/doc/user/Delegation.md +191 -0
  19. data/doc/user/GettingStarted.md +159 -0
  20. data/doc/user/ManagingContext.md +163 -0
  21. data/doc/user/MultiAgentWorkflows.md +256 -0
  22. data/doc/user/Python.md +159 -0
  23. data/doc/user/RunningInference.md +200 -0
  24. data/doc/user/ToolCalling.md +193 -0
  25. data/doc/user/WritingChats.md +197 -0
  26. data/lib/scout/llm/agent/chat.rb +61 -11
  27. data/lib/scout/llm/agent/delegate.rb +274 -65
  28. data/lib/scout/llm/agent/iterate.rb +2 -2
  29. data/lib/scout/llm/agent/save.rb +273 -0
  30. data/lib/scout/llm/agent/workflow.rb +164 -0
  31. data/lib/scout/llm/agent.rb +86 -61
  32. data/lib/scout/llm/ask.rb +62 -17
  33. data/lib/scout/llm/backends/anthropic.rb +9 -2
  34. data/lib/scout/llm/backends/bedrock.rb +15 -3
  35. data/lib/scout/llm/backends/default.rb +183 -99
  36. data/lib/scout/llm/backends/glm.rb +58 -0
  37. data/lib/scout/llm/backends/huggingface.rb +196 -26
  38. data/lib/scout/llm/backends/ollama.rb +13 -1
  39. data/lib/scout/llm/backends/openai.rb +0 -2
  40. data/lib/scout/llm/backends/openwebui.rb +20 -13
  41. data/lib/scout/llm/backends/relay.rb +22 -22
  42. data/lib/scout/llm/backends/responses.rb +1 -1
  43. data/lib/scout/llm/chat/agent_meta.rb +264 -0
  44. data/lib/scout/llm/chat/annotation.rb +39 -10
  45. data/lib/scout/llm/chat/parse.rb +28 -6
  46. data/lib/scout/llm/chat/persist.rb +25 -0
  47. data/lib/scout/llm/chat/process/clear.rb +41 -6
  48. data/lib/scout/llm/chat/process/files.rb +21 -6
  49. data/lib/scout/llm/chat/process/meta.rb +421 -34
  50. data/lib/scout/llm/chat/process/options.rb +21 -1
  51. data/lib/scout/llm/chat/process/tools.rb +56 -15
  52. data/lib/scout/llm/chat/process.rb +4 -0
  53. data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
  54. data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
  55. data/lib/scout/llm/chat/prompt.rb +48 -0
  56. data/lib/scout/llm/chat/provenance.rb +775 -0
  57. data/lib/scout/llm/chat/tool_calls.rb +76 -0
  58. data/lib/scout/llm/chat.rb +18 -2
  59. data/lib/scout/llm/embed.rb +11 -3
  60. data/lib/scout/llm/image.rb +86 -0
  61. data/lib/scout/llm/mcp.rb +10 -2
  62. data/lib/scout/llm/rag.rb +3 -3
  63. data/lib/scout/llm/tools/call.rb +160 -11
  64. data/lib/scout/llm/tools/knowledge_base.rb +1 -1
  65. data/lib/scout/llm/tools/workflow.rb +32 -16
  66. data/lib/scout/model/python/huggingface/causal.rb +23 -5
  67. data/lib/scout/model/python/huggingface.rb +2 -1
  68. data/lib/scout-ai.rb +1 -0
  69. data/python/README.md +197 -14
  70. data/python/scout_ai/huggingface/eval.py +245 -34
  71. data/python/tests/test_huggingface_eval.py +58 -0
  72. data/research/ChatAnalyst-required-changes.md +167 -0
  73. data/research/agent-delegation-analysis.md +810 -0
  74. data/research/agent-meta-provenance-integration-plan.md +622 -0
  75. data/research/agent-workflow-analysis.md +1120 -0
  76. data/research/backends-analysis.md +836 -0
  77. data/research/chat-core-analysis.md +946 -0
  78. data/research/chatanalyst-provenance/00-baseline.md +30 -0
  79. data/research/chatanalyst-provenance/01-repo-map.md +60 -0
  80. data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
  81. data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
  82. data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
  83. data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
  84. data/research/chatanalyst-provenance/07-critic-review.md +25 -0
  85. data/research/chatanalyst-provenance/final-report.md +45 -0
  86. data/research/chatanalyst-provenance/resumption.md +37 -0
  87. data/research/coding-philosophy-analysis.md +928 -0
  88. data/research/commands-analysis.md +947 -0
  89. data/research/multi-agent-patterns-analysis.md +853 -0
  90. data/research/prompt-strategies-analysis.md +630 -0
  91. data/research/prov-verbosity-fix-notes.md +77 -0
  92. data/research/provenance-analysis.md +469 -0
  93. data/research/provenance-navigation-design.md +640 -0
  94. data/research/synthesis-report.md +487 -0
  95. data/research/tools-system-analysis.md +779 -0
  96. data/scout-ai.gemspec +100 -11
  97. data/scout_commands/agent/ask +13 -3
  98. data/scout_commands/agent/kb +2 -0
  99. data/scout_commands/llm/ask +11 -4
  100. data/scout_commands/llm/md +76 -0
  101. data/scout_commands/llm/process_queries +48 -0
  102. data/scout_commands/llm/prov +602 -0
  103. data/scout_commands/llm/word +71 -0
  104. data/scout_commands/workflow/mcp +43 -0
  105. data/share/word/reference.docx +0 -0
  106. data/test/etc/AI/mock.yaml +11 -0
  107. data/test/fixtures/backends/anthropic.json +19 -0
  108. data/test/fixtures/backends/anthropic_tool_use.json +24 -0
  109. data/test/fixtures/backends/bedrock.json +8 -0
  110. data/test/fixtures/backends/bedrock_embedding.json +3 -0
  111. data/test/fixtures/backends/bedrock_tool_use.json +17 -0
  112. data/test/fixtures/backends/ollama.json +16 -0
  113. data/test/fixtures/backends/ollama_tool_call.json +27 -0
  114. data/test/fixtures/backends/openai_chat.json +21 -0
  115. data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
  116. data/test/fixtures/backends/responses.json +33 -0
  117. data/test/fixtures/backends/responses_tool_call.json +28 -0
  118. data/test/integration/README.md +32 -0
  119. data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
  120. data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
  121. data/test/integration/scout/llm/backends/test_relay.rb +52 -0
  122. data/test/integration/scout/llm/test_infrastructure.rb +74 -0
  123. data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
  124. data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
  125. data/test/integration/scout/model/test_base.rb +91 -0
  126. data/test/scout/llm/agent/test_chat.rb +8 -2
  127. data/test/scout/llm/agent/test_save.rb +413 -0
  128. data/test/scout/llm/agent/test_workflow.rb +110 -0
  129. data/test/scout/llm/backends/test_anthropic.rb +93 -10
  130. data/test/scout/llm/backends/test_bedrock.rb +118 -2
  131. data/test/scout/llm/backends/test_huggingface.rb +137 -42
  132. data/test/scout/llm/backends/test_ollama.rb +70 -20
  133. data/test/scout/llm/backends/test_openwebui.rb +42 -40
  134. data/test/scout/llm/backends/test_relay.rb +4 -2
  135. data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
  136. data/test/scout/llm/chat/process/test_meta.rb +518 -0
  137. data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
  138. data/test/scout/llm/chat/test_agent_meta.rb +357 -0
  139. data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
  140. data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
  141. data/test/scout/llm/chat/test_parse.rb +70 -15
  142. data/test/scout/llm/chat/test_prov_cli.rb +274 -0
  143. data/test/scout/llm/chat/test_provenance.rb +240 -0
  144. data/test/scout/llm/chat/test_tool_calls.rb +38 -0
  145. data/test/scout/llm/test_agent.rb +13 -36
  146. data/test/scout/llm/test_ask.rb +75 -52
  147. data/test/scout/llm/test_chat.rb +107 -13
  148. data/test/scout/llm/test_embed.rb +48 -0
  149. data/test/scout/llm/test_rag.rb +23 -16
  150. data/test/scout/llm/test_tools.rb +12 -1
  151. data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
  152. data/test/scout/llm/tools/test_mcp.rb +5 -3
  153. data/test/scout/llm/tools/test_workflow.rb +23 -2
  154. data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
  155. data/test/scout/model/python/huggingface/test_causal.rb +9 -3
  156. data/test/scout/model/python/huggingface/test_classification.rb +11 -2
  157. data/test/scout/model/python/test_torch.rb +2 -0
  158. data/test/scout/model/python/torch/test_helpers.rb +4 -0
  159. data/test/scout/model/test_base.rb +4 -2
  160. data/test/support/availability.rb +231 -0
  161. data/test/support/fake_clients.rb +138 -0
  162. data/test/support/fixtures.rb +21 -0
  163. data/test/support/infrastructure_probes.rb +136 -0
  164. data/test/support/mock_backend.rb +215 -0
  165. data/test/test_helper.rb +32 -2
  166. metadata +99 -10
  167. data/doc/Agent.md +0 -327
  168. data/doc/Chat.md +0 -458
  169. data/doc/LLM.md +0 -340
  170. data/doc/RAG.md +0 -129
  171. data/scout_commands/documenter +0 -148
  172. data/test/scout/llm/backends/test_openai.rb +0 -192
  173. data/test/scout/llm/backends/test_responses.rb +0 -238
  174. data/test/scout/llm/test_parse.rb +0 -98
@@ -0,0 +1,77 @@
1
+ # prov verbosity fix notes
2
+
3
+ ## Problem
4
+
5
+ The `scout-ai llm prov` command output was too verbose and token counts were
6
+ wrong.
7
+
8
+ Two issues:
9
+
10
+ 1. **Verbosity**: The output showed many redundant nodes — `agent.chat` log
11
+ files, zero-token result chats, `(seen)` repeated nodes, inline relation
12
+ labels, and long filesystem paths.
13
+
14
+ 2. **Token counts**: The root chat showed `total=0` because it had no direct
15
+ inference tokens. Every parent node should show the aggregate cost of its
16
+ entire provenance subtree.
17
+
18
+ ## Root causes
19
+
20
+ ### Job resolution
21
+
22
+ `Step.load` with a relative workflow reference like
23
+ `Planned/ask/Default_xyz.chat` resolved via `Path.find` to
24
+ `~/.scout/Planned/ask/Default_xyz.chat`, but the actual job data lives at
25
+ `~/.rbbt/var/jobs/Planned/ask/Default_xyz.chat`. The `.scout` directory did not
26
+ contain the job files, so traversal stopped immediately.
27
+
28
+ ### Redundant nodes
29
+
30
+ The traversal visits chat files connected to jobs via four relations: `job`,
31
+ `dependency`, `log`, `result`. Of these:
32
+
33
+ - **`log` relation with `agent.chat`**: This is the chat produced by the
34
+ inference backend, stored at `<job>.files/log/agent.chat`. Its tokens are the
35
+ same tokens counted as the job's direct cost. Showing it as a separate node
36
+ duplicates the parent's cost.
37
+ - **`result` relation**: A chat-typed job's persisted result file IS the same
38
+ path as the job itself. The `[:chat, path]` node is a duplicate of the
39
+ `[:job, path]` node.
40
+ - **Repeated nodes in shared DAG branches**: When a dependency appears in
41
+ multiple branches, it was shown as `(seen)` in the old verbose output.
42
+
43
+ ### Token aggregation
44
+
45
+ The old code only showed each node's **direct** token cost (from its own
46
+ `agent.chat` or log files). This meant a root chat with no direct inference
47
+ showed `total=0`. The correct behavior is to aggregate the entire subtree's
48
+ inference cost.
49
+
50
+ ## Fixes applied
51
+
52
+ ### `lib/scout/llm/chat/provenance.rb`
53
+
54
+ Added `Chat.load_job_reference(reference)` helper that tries the standard
55
+ `~/.rbbt/var/jobs/` storage location as a fallback when `Step.load` resolves to
56
+ a non-existent path. The traversal uses this for all `chat.jobs` references.
57
+
58
+ ### `scout_commands/llm/prov`
59
+
60
+ 1. **Hidden nodes**: `agent.chat` log files and result-relation chats are
61
+ skipped entirely in tree mode (not printed, not traversed into).
62
+
63
+ 2. **No `(seen)` lines**: Repeated nodes are silently skipped in tree mode.
64
+
65
+ 3. **No inline relation labels**: Tree lines show just `kind tokens label`.
66
+
67
+ 4. **Shortened paths**: Log chats under jobs show their path relative to the
68
+ job's `.files/log/` directory (e.g. `society/Worker/cli_investigation.chat`).
69
+
70
+ 5. **Aggregate mode** (default): Each node shows the total cost of all
71
+ reachable chat files in its provenance subtree, deduplicated by path.
72
+
73
+ 6. **Component mode** (`--component`): Each node shows only its own direct
74
+ token cost.
75
+
76
+ 7. **Flow and DOT modes**: Same hidden-node filtering applied. Flow mode shows
77
+ data-flow arrows (result/dependency/log) between visible nodes.
@@ -0,0 +1,469 @@
1
+ > **Disclaimer:** This is an architectural investigation, not normative
2
+ > documentation. It was produced during a documentation-revamp effort and may
3
+ > be outdated relative to the current codebase. Treat it as supporting
4
+ > reference material. For maintained documentation, see
5
+ > [../../doc/](../../doc/).
6
+ >
7
+
8
+
9
+ # Provenance System: Chat.provenance, trace_chats, prov/info Commands, and ChatAnalyst
10
+
11
+ > **Scope:** How Scout-AI records, traverses, and reports the lineage of every
12
+ > inference — across imported chats, ask-job results, nested agent logs, and
13
+ > dependency chains — and the two CLI commands (`prov` and `info`) and the
14
+ > ChatAnalyst workflow that consume that provenance data.
15
+
16
+ ---
17
+
18
+ ## Provenance Data Model
19
+
20
+ ### What provenance metadata is stored
21
+
22
+ Scout-AI does not have a standalone "provenance" table. Instead, provenance
23
+ information is embedded in **meta messages** that appear inline within the
24
+ chat transcript itself. A meta message has `role: meta` and a content string
25
+ serialized as `key=value key=value ...`.
26
+
27
+ The serialization/parsing layer lives in
28
+ `lib/scout/llm/chat/process/meta.rb`:
29
+
30
+ | Method | Purpose |
31
+ |---|---|
32
+ | `Chat.serialize_meta(hash)` | Converts a hash to the `key=value` string format, sorted by value length. |
33
+ | `Chat.parse_meta(str)` | Parses a `key=value` string back into an `IndiferentHash`. |
34
+ | `Chat.meta(messages)` | Strips all `meta` messages from an array and returns a merged metadata hash. The last checkpoint with `pt_c`/`ct_c`/`tt_c` fields is used for cumulative counters. |
35
+
36
+ #### Fields that can appear in a meta message
37
+
38
+ **Token fields** (written by `Backend::Default#update_meta` in
39
+ `lib/scout/llm/backends/default.rb`):
40
+
41
+ | Field | Meaning |
42
+ |---|---|
43
+ | `pt` | Prompt tokens for this single inference |
44
+ | `ct` | Completion tokens for this single inference |
45
+ | `tt` | Total tokens for this single inference |
46
+ | `pt_s`, `ct_s`, `tt_s` | **Session** cumulative counters (per-thread running totals) |
47
+ | `pt_c`, `ct_c`, `tt_c` | **Chat** cumulative counters (persisted across requests) |
48
+ | `reas` | Reasoning summary string (truncated for display) |
49
+
50
+ **Job-reference field** (written when a chat-task result is projected):
51
+
52
+ | Field | Meaning |
53
+ |---|---|
54
+ | `job` | The canonical path of the Scout workflow job that produced this segment |
55
+
56
+ #### Two kinds of meta messages
57
+
58
+ 1. **Direct inference meta** — contains `pt`, `ct`, `tt` (and optionally
59
+ cumulative/session variants and `reas`). This records one actual model call.
60
+
61
+ 2. **Job projection meta** — contains only `job=<path>`. This marks a response
62
+ segment that was projected from an ask-workflow job. It has **zero direct
63
+ token cost** — the actual tokens are recorded in the job's own agent logs
64
+ and dependency chain.
65
+
66
+ ### How annotation.rb tracks provenance
67
+
68
+ The file `lib/scout/llm/chat/annotation.rb` provides the `Chat` annotation
69
+ (through `extend Annotation`). It adds convenience methods like `user`,
70
+ `assistant`, `import`, `option`, `endpoint`, `model`, etc. But it does **not**
71
+ contain provenance-specific methods — those all live in
72
+ `lib/scout/llm/chat/process/meta.rb`.
73
+
74
+ Relevant provenance-related instance methods on `Chat` objects (all defined in
75
+ `process/meta.rb`):
76
+
77
+ | Method | Returns |
78
+ |---|---|
79
+ | `chat.job_paths` / `chat.jobs` | Array of `Path` objects extracted from all `meta job=...` messages |
80
+ | `chat.job_chat_files` | All chat files (result + logs) reachable from this chat's jobs and their dependencies |
81
+ | `chat.job_agent_chat_files` | All `log/**/*.chat` files from this chat's jobs and their dependencies |
82
+ | `chat.job_chats` | All `Chat` objects loaded from `job_chat_files` |
83
+ | `chat.job_agent_chats` | All `Chat` objects loaded from `job_agent_chat_files` |
84
+ | `chat.message_index` | Array of per-message lineage records with `id`, `role`, `prev`, `fingerprint`, and parsed `meta` |
85
+ | `chat.meta` | The parsed metadata hash from the last meta message |
86
+ | `chat.last_job` | The `job` value from the last meta message |
87
+
88
+ ### The `Chat.project` method
89
+
90
+ When a chat-task produces output, `Chat.project(job, messages)` wraps the
91
+ non-meta messages with a single `meta job=<path>` marker at the front. This
92
+ ensures that consumers can detect the job origin without scanning for token
93
+ fields:
94
+
95
+ ```ruby
96
+ [{ role: :meta, content: serialize_meta(job: job.to_s) }] + projected_messages
97
+ ```
98
+
99
+ ---
100
+
101
+ ## Recursive Traversal / trace_chats
102
+
103
+ ### Lineage IDs and message_index
104
+
105
+ `Chat#message_index` (defined in `process/meta.rb`) computes a **lineage ID**
106
+ for each message:
107
+
108
+ ```ruby
109
+ id = Misc.digest([previous_id, role, content])
110
+ ```
111
+
112
+ Each message's lineage ID incorporates the previous conversational message's ID
113
+ (meta messages are excluded from the lineage chain — they start segments but
114
+ are not provider input). This creates a hash-chain where `prev` links each
115
+ message to its predecessor in the *conversational* history.
116
+
117
+ The index also assigns each message a `fingerprint` (a truncated head/tail
118
+ digest via `Log.truncate_string`) for compact comparison.
119
+
120
+ ### Response segments and trace_indices
121
+
122
+ `Chat.trace_indices(indices)` walks a set of message indices and groups them
123
+ into **response segments**. The algorithm:
124
+
125
+ 1. Iterate through messages in order.
126
+ 2. When a `meta` message is encountered, **close** the current pending segment
127
+ (if any) and **open** a new one, seeded with the parsed metadata.
128
+ 3. `user` and `system` messages also close any pending segment.
129
+ 4. All other messages (`assistant`, `function_call`, `function_call_output`,
130
+ `tool`, etc.) are appended to the current segment's message list.
131
+ 5. At the end, close any remaining pending segment.
132
+ 6. Each segment gets `orphan: true` if it has zero covered messages.
133
+
134
+ A `seen` Set of lineage IDs prevents double-counting.
135
+
136
+ ### trace_chats
137
+
138
+ ```ruby
139
+ def self.trace_chats(chats)
140
+ trace_indices(chats.collect(&:message_index))
141
+ end
142
+ ```
143
+
144
+ This takes an array of `Chat` objects, computes `message_index` on each, then
145
+ runs `trace_indices` across all of them. The result is a flat array of segment
146
+ records:
147
+
148
+ ```ruby
149
+ { id: <lineage_id>, meta: <parsed_meta_hash>, messages: [<id>, ...], orphan: true|false }
150
+ ```
151
+
152
+ ### What trace_chats returns
153
+
154
+ Each entry represents one **response segment** — one model call (or one job
155
+ projection). The `meta` field tells you whether it's a direct inference (has
156
+ `pt`/`ct`/`tt`) or a projection (has `job`). The `messages` array lists the
157
+ lineage IDs of all messages that belong to that segment.
158
+
159
+ **Token accounting from the trace:** To count tokens, filter to entries where
160
+ `meta` has no `:job` key but has `pt`/`ct`/`tt`, then sum those fields. The
161
+ `*_c` and `*_s` counters are checkpoints and must never be summed — they would
162
+ double-count.
163
+
164
+ ### Job-based recursive provenance
165
+
166
+ The `Chat.job_chat_files(job, seen)` class method performs recursive traversal
167
+ of the **job dependency graph**:
168
+
169
+ 1. Load the job via `Step.load`.
170
+ 2. If the job is done and its type is `chat`, add its result path.
171
+ 3. Add all `log/**/*.chat` files from the job's `files_dir`.
172
+ 4. For each dependency, recurse (using a `seen` Set to prevent cycles and
173
+ duplicate visits).
174
+ 5. Return the unique set of all discovered chat file paths.
175
+
176
+ This is the mechanism that `chat.job_chat_files` (instance method) uses to
177
+ discover the full provenance tree below a chat.
178
+
179
+ ---
180
+
181
+ ## The `prov` Command
182
+
183
+ **File:** `scout_commands/llm/prov`
184
+
185
+ ### What it does
186
+
187
+ `prov` prints a hierarchical, indented tree showing the provenance structure
188
+ below a chat file or job path. It walks jobs → agent logs → dependencies
189
+ recursively and prints token totals at each node.
190
+
191
+ ### Usage
192
+
193
+ ```bash
194
+ scout-ai llm prov <filename>
195
+ ```
196
+
197
+ ### How it works (internals)
198
+
199
+ The `prov` command **monkey-patches** several methods onto the `Chat` and
200
+ `Step` classes at runtime (these are NOT in the library):
201
+
202
+ - `Chat.provenance(chat_file, prov={})` — recursively walks a chat's jobs,
203
+ their agent chat files, and dependencies, building a hash mapping each chat
204
+ file to the list of agent chat files it references.
205
+ - `Chat.provenance_chat_files(chat)` — flattens the provenance hash into a
206
+ unique list of all chat files.
207
+ - `Chat.tokens(chat)` — loads all provenance chat files, then sums `pt`/`ct`/
208
+ `tt` from direct entries only.
209
+ - `Chat.trace`, `Chat.direct_entries`, `Chat.token_totals`,
210
+ `Chat.print_tokens` — helper methods for trace-based token accounting.
211
+ - `Chat.job_agent_chat_files(job)` — **redefines** the library method.
212
+ - `Step#agent_chats` — new method on Step.
213
+
214
+ These monkey-patches mean `prov` is self-contained but diverges from the
215
+ library's actual API. The `Chat.job_agent_chat_files` in the library already
216
+ exists and works similarly; the `prov` version is a redundant redefinition.
217
+
218
+ ### Output format
219
+
220
+ ```
221
+ job total=1.2k prompt=800 cont=400 ~/.scout/var/jobs/.../ask
222
+ chat total=500 prompt=300 cont=200 agent.chat
223
+ job total=700 prompt=500 cont=200 ~/.scout/var/jobs/.../subtask
224
+ ```
225
+
226
+ - **Yellow `job`** lines show job paths with token totals.
227
+ - **Green `chat`** lines show chat/agent-log files with token totals.
228
+ - Indentation reflects the nesting depth.
229
+ - Token totals are computed by summing `pt`/`ct`/`tt` from direct inference
230
+ metas across all provenance chat files.
231
+
232
+ ### Limitations
233
+
234
+ - The `agent.chat` file (the default agent log) is suppressed in the output
235
+ (`unless name == 'agent.chat'`), but its tokens are still counted.
236
+ - The command has a hardcoded fallback filename:
237
+ `~/git/workflows/SC26/chats/network_usecase/3.1.themes` if no argument is
238
+ given.
239
+ - It does not produce machine-readable output (no JSON mode).
240
+
241
+ ---
242
+
243
+ ## The `info` Command — STATUS ASSESSMENT
244
+
245
+ **File:** `scout_commands/llm/info`
246
+
247
+ ### What it does
248
+
249
+ `info` is a **substantially more capable** command than `prov`. It:
250
+
251
+ 1. **Discovers the full provenance graph** — root chat, imports, job results,
252
+ job dependencies, and agent logs — using a BFS traversal.
253
+ 2. **Deduplicates jobs** by a canonical identity (workflow/task/basename) so
254
+ that `~/.scout` and `~/.rbbt` mirrors of the same job appear once.
255
+ 3. **Reports** chats (with role/message counts and job references), jobs (with
256
+ workflow/task, log count, dependency count), and token usage (root-only vs
257
+ all-traced).
258
+ 4. **Optional flow mode** (`-f`/`--flow`): prints a compact numbered node/edge
259
+ list with token annotations.
260
+ 5. **Optional DOT/plot mode** (`--dot`, `--plot`): generates Graphviz DOT or
261
+ renders SVG/PNG/PDF.
262
+
263
+ ### Is `info` outdated? — **NO, it is current and well-maintained**
264
+
265
+ The user suspected `info` might be outdated. After thorough cross-referencing,
266
+ **the `info` command is NOT outdated**. It is in fact the more modern and
267
+ complete implementation. Here is the specific evidence:
268
+
269
+ #### Methods/APIs it calls and their current status
270
+
271
+ | API used by `info` | Defined in | Still exists? |
272
+ |---|---|---|
273
+ | `Chat.load(path)` | `lib/scout/llm/chat/process/meta.rb:86` | ✅ Yes |
274
+ | `Chat.trace_chats(chats)` | `lib/scout/llm/chat/process/meta.rb:195` | ✅ Yes |
275
+ | `Chat.find_file(...)` | `lib/scout/llm/chat/process/files.rb:17` | ✅ Yes |
276
+ | `chat.role_messages(role)` | `lib/scout/llm/chat/annotation.rb` | ✅ Yes |
277
+ | `chat.jobs` / `chat.job_paths` | `lib/scout/llm/chat/process/meta.rb:76` | ✅ Yes |
278
+ | `job.dependencies` | Scout `Step` API | ✅ Yes |
279
+ | `job.file('log')` | Scout `Step` API | ✅ Yes |
280
+ | `job.info[:workflow]`, `job.info[:task_name]` | Scout `Step` API | ✅ Yes |
281
+ | `job.done?`, `job.type` | Scout `Step` API | ✅ Yes |
282
+ | `Step.load(path)` | Scout API | ✅ Yes |
283
+
284
+ #### Annotation fields it reads
285
+
286
+ The `info` command reads `pt`, `ct`, `tt` from parsed meta messages via
287
+ `Chat.trace_chats` → `trace_indices` → per-entry `:meta` hash. These fields are
288
+ written by `Backend::Default#update_meta` (line 420 of `backends/default.rb`)
289
+ and are fully current.
290
+
291
+ It correctly filters to **direct entries only** (entries whose meta has no
292
+ `:job` key but has `pt`/`ct`/`tt`), matching the exact pattern used by
293
+ `ChatAnalyst::Session#token_entries`.
294
+
295
+ #### What `info` does that `prov` does not
296
+
297
+ | Feature | `info` | `prov` |
298
+ |---|---|---|
299
+ | Import discovery | ✅ (`import`/`continue`/`last` roles) | ❌ |
300
+ | Job deduplication | ✅ (canonical identity by workflow/task/basename) | ❌ |
301
+ | Dependency graph | ✅ | ✅ (via Chat.provenance) |
302
+ | Flow visualization | ✅ (text flow + Graphviz DOT/SVG/PNG/PDF) | ❌ |
303
+ | Warnings | ✅ (records load failures) | ❌ |
304
+ | Token accounting | ✅ (`trace_chats` + direct_entries) | ✅ (same approach, but monkey-patched) |
305
+ | Uses library API directly | ✅ (no monkey-patching) | ❌ (redefines Chat methods) |
306
+
307
+ #### Assessment summary
308
+
309
+ - **`info` is the current, recommended command** for provenance inspection.
310
+ - **`prov` is older** and relies on runtime monkey-patches rather than the
311
+ library API. It still works but is superseded by `info` in functionality.
312
+ - Neither command is "broken" — both will execute successfully.
313
+ - If anything is "outdated," it is `prov` (monkey-patches, no import
314
+ discovery, no flow/DOT output), not `info`.
315
+
316
+ ---
317
+
318
+ ## ChatAnalyst Provenance Capabilities
319
+
320
+ **Location:** `~/git/workflows/SC26/Agent/ChatAnalyst/`
321
+
322
+ ### What the ChatAnalyst agent does
323
+
324
+ ChatAnalyst is a Scout-AI agent specialized in **inspecting work sessions** to
325
+ find usage patterns, tool-calling issues, barriers, and areas for improving
326
+ the harness, agent instructions, or tooling.
327
+
328
+ Its system prompt (`start_chat`) directs it to:
329
+ - Follow conversations across different chats and ask jobs.
330
+ - Inspect persisted chat sessions for patterns and issues.
331
+ - Use the ChatAnalyst workflow tooling (README.md) for structured analysis.
332
+
333
+ ### The Session class (workflow.rb core)
334
+
335
+ The `workflow.rb` defines a `ChatAnalyst::Session` class that performs
336
+ BFS-based discovery identical in spirit to `info`'s `LLMInfoReport`:
337
+
338
+ 1. **`resolve_chat(input)`** — tries multiple candidate paths (plain,
339
+ `.chat` extension, expanded, `Scout.chats[]` lookup).
340
+
341
+ 2. **`discover_chat(path)`** — for each chat:
342
+ - Records it in `@chats`.
343
+ - Follows `import`/`continue`/`last` roles to discover imported chats
344
+ (adding `import` edges).
345
+ - Follows `meta job=...` references to discover producer jobs (adding
346
+ `result` edges).
347
+
348
+ 3. **`discover_job(reference)`** — for each job:
349
+ - Records it in `@jobs`.
350
+ - Follows dependencies recursively (adding `dependency` edges).
351
+ - If the job is done and type is `chat`, discovers the result chat (adding
352
+ it to the chat queue).
353
+ - Scans `log/**/*.chat` for agent conversation logs (adding `log` edges and
354
+ feeding them back into `discover_chat`).
355
+
356
+ 4. **Token accounting** — `token_entries` filters the trace to direct inference
357
+ segments (no `:job` key, has `pt`/`ct`/`tt`). `token_totals` sums them.
358
+
359
+ ### Tasks exposed by the ChatAnalyst workflow
360
+
361
+ | Task | Type | Purpose |
362
+ |---|---|---|
363
+ | `message_index` | JSON | Compact per-message index across the full session tree. Each entry has an ID (`file#index`), lineage ID, previous lineage ID, role, fingerprint, and parsed meta. Supports optional `role` filter. |
364
+ | `message_content` | JSON | Retrieves full untruncated content for specific message IDs. Two-phase inspection: scan index → drill into interesting messages. |
365
+ | `chat_overview` | JSON | Structural graph overview: per-file summaries (roles, messages, jobs, tool calls), per-job summaries (workflow/task, dependencies, logs), typed edges, aggregate totals, warnings. |
366
+ | `chat_tool_calls` | JSON | Pairs `function_call`/`mcp_call` with `function_call_output` by call ID. Reports tool name, call ID, output position, success/failure, exception/exit status. Summary: totals, successes, failures, breakdown by tool. |
367
+ | `chat_tokens` | JSON | Per-file and aggregate token accounting using direct `pt`/`ct`/`tt` metas only. Includes `note` explaining methodology. |
368
+ | `chat_agents` | JSON | Detects `ask` and `hand_off_to_*` calls. Reports interactions with source file, call ID, output position, success state. |
369
+ | `chat_report` | JSON | Combined snapshot: session size, job count, aggregate tokens, trace records, first 20 tool calls, all failures, warnings. |
370
+
371
+ ### Tooling file
372
+
373
+ The `tooling` file in the ChatAnalyst agent directory is a **chat session
374
+ log** (not a tooling declaration file). It contains the conversation in which
375
+ the ChatAnalyst workflow was originally designed, including a detailed
376
+ critique and README update session. It is not loaded as tooling — it is a
377
+ record of the development process.
378
+
379
+ ### How ChatAnalyst compares to `info`
380
+
381
+ | Aspect | `info` CLI | ChatAnalyst workflow |
382
+ |---|---|---|
383
+ | Consumer | Human (terminal output, DOT/SVG) | Agent (JSON tasks) |
384
+ | Discovery | Identical BFS (imports, jobs, deps, logs) | Identical BFS |
385
+ | Token accounting | `trace_chats` + direct entries | `trace_chats` + direct entries |
386
+ | Job deduplication | ✅ (canonical identity) | ❌ (simple expand_path check) |
387
+ | Tool-call analysis | ❌ | ✅ |
388
+ | Agent-interaction analysis | ❌ | ✅ |
389
+ | Output | Text / Graphviz | JSON |
390
+ | Graphviz rendering | ✅ | ❌ |
391
+
392
+ ---
393
+
394
+ ## Key Design Notes
395
+
396
+ ### The provenance tree metaphor
397
+
398
+ A Scout-AI session is not a single flat conversation. It is a **tree** (or DAG)
399
+ of conversations connected by:
400
+
401
+ 1. **Import edges** — `import:`/`continue:`/`last:` roles pull in previous
402
+ chat history, creating a horizontal chain of related conversations.
403
+
404
+ 2. **Result edges** — `meta: job=<path>` markers indicate that a response
405
+ segment was produced by a Scout workflow job (typically an `ask` task).
406
+ The actual inference happened inside that job's agent logs.
407
+
408
+ 3. **Dependency edges** — Scout workflow jobs have dependencies (other jobs
409
+ they consume). Each dependency may itself be a chat-producing job with its
410
+ own agent logs.
411
+
412
+ 4. **Log edges** — Each ask-job has a `log/` directory containing agent chat
413
+ files. The primary `log/agent.chat` is the agent's full conversation.
414
+ Socialized projections may appear under `log/chats/<AgentName>/`.
415
+
416
+ 5. **Call edges** (semantic, not structural) — Tool calls to `ask` or
417
+ `hand_off_to_*` represent agent-to-agent delegation. These are detected by
418
+ scanning tool-call messages, not by following meta references.
419
+
420
+ ### How multi-agent inference creates provenance chains
421
+
422
+ When an orchestrator agent (e.g., `AGI`) dispatches work to a specialist agent
423
+ (e.g., `Worker`) via `ask`:
424
+
425
+ 1. The orchestrator's chat contains a `function_call` with tool name `ask` and
426
+ arguments naming the target agent.
427
+ 2. The `ask` call triggers a Scout workflow job (e.g.,
428
+ `Agent/Worker/ask`).
429
+ 3. That job runs the target agent, producing `log/agent.chat` with the full
430
+ agent conversation and its own `meta:` entries with direct token counts.
431
+ 4. The job result is projected back into the orchestrator's chat as a
432
+ `meta: job=<path>` marker followed by the response messages.
433
+ 5. The orchestrator's chat thus has a **zero-token projection segment** — the
434
+ real tokens are in the job's `agent.chat`.
435
+
436
+ This means token accounting requires **recursive traversal**: start at the root
437
+ chat, follow every `meta: job=...` to its job, read the job's `agent.chat`
438
+ logs, follow the job's dependencies, and sum only direct `pt`/`ct`/`tt`
439
+ entries. Cumulative (`*_c`) and session (`*_s`) counters must never be summed
440
+ because they would double-count across the import/projection chain.
441
+
442
+ ### Socialized chat projections
443
+
444
+ When an agent dispatches to another agent with a named `conversation`, the
445
+ specialist interaction may be persisted as a socialized chat at:
446
+ ```
447
+ <caller_job>.files/log/chats/<AgentName>/<conversation>.chat
448
+ ```
449
+ These files contain the prompt, propagated `option:` lines, a `meta: job=...`
450
+ marker, and the response — but **zero direct inference tokens**. The real
451
+ model calls are found by following the `meta: job=...` reference into the
452
+ specialist's ask-job and its `agent.chat` log.
453
+
454
+ ### Why `info` is preferred over `prov`
455
+
456
+ - `info` uses the library API directly (no monkey-patching).
457
+ - `info` discovers imports (prov does not).
458
+ - `info` deduplicates mirrored jobs between `~/.scout` and `~/.rbbt`.
459
+ - `info` supports flow visualization (text + Graphviz).
460
+ - `info` records and reports warnings for load failures.
461
+ - `prov` remains functional but is architecturally older and less complete.
462
+
463
+ ### Design consistency between `info` and ChatAnalyst
464
+
465
+ Both `info` and ChatAnalyst use the same `Chat.trace_chats` API for token
466
+ accounting and the same BFS pattern for provenance discovery. The key
467
+ difference is that `info` adds job deduplication (canonical identity), while
468
+ ChatAnalyst adds tool-call and agent-interaction analysis. They are
469
+ complementary tools for the same provenance data model.