scout-ai 1.2.3 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.vimproject +138 -50
- data/README.md +171 -290
- data/Rakefile +17 -1
- data/VERSION +1 -1
- data/doc/Improvements.md +325 -0
- data/doc/StartHere.md +110 -0
- data/doc/developer/Architecture.md +126 -0
- data/doc/developer/Backends.md +199 -0
- data/doc/developer/ChatLifecycle.md +183 -0
- data/doc/developer/DelegationInternals.md +295 -0
- data/doc/developer/DesignPrinciples.md +245 -0
- data/doc/developer/PromptProcessing.md +292 -0
- data/doc/developer/Provenance.md +317 -0
- data/doc/user/BuildingAgents.md +345 -0
- data/doc/user/Cookbook.md +333 -0
- data/doc/user/CoreConcepts.md +181 -0
- data/doc/user/Delegation.md +191 -0
- data/doc/user/GettingStarted.md +159 -0
- data/doc/user/ManagingContext.md +163 -0
- data/doc/user/MultiAgentWorkflows.md +256 -0
- data/doc/user/Python.md +159 -0
- data/doc/user/RunningInference.md +200 -0
- data/doc/user/ToolCalling.md +193 -0
- data/doc/user/WritingChats.md +197 -0
- data/lib/scout/llm/agent/chat.rb +61 -11
- data/lib/scout/llm/agent/delegate.rb +274 -65
- data/lib/scout/llm/agent/iterate.rb +2 -2
- data/lib/scout/llm/agent/save.rb +273 -0
- data/lib/scout/llm/agent/workflow.rb +164 -0
- data/lib/scout/llm/agent.rb +86 -61
- data/lib/scout/llm/ask.rb +62 -17
- data/lib/scout/llm/backends/anthropic.rb +9 -2
- data/lib/scout/llm/backends/bedrock.rb +15 -3
- data/lib/scout/llm/backends/default.rb +183 -99
- data/lib/scout/llm/backends/glm.rb +58 -0
- data/lib/scout/llm/backends/huggingface.rb +196 -26
- data/lib/scout/llm/backends/ollama.rb +13 -1
- data/lib/scout/llm/backends/openai.rb +0 -2
- data/lib/scout/llm/backends/openwebui.rb +20 -13
- data/lib/scout/llm/backends/relay.rb +22 -22
- data/lib/scout/llm/backends/responses.rb +1 -1
- data/lib/scout/llm/chat/agent_meta.rb +264 -0
- data/lib/scout/llm/chat/annotation.rb +39 -10
- data/lib/scout/llm/chat/parse.rb +28 -6
- data/lib/scout/llm/chat/persist.rb +25 -0
- data/lib/scout/llm/chat/process/clear.rb +41 -6
- data/lib/scout/llm/chat/process/files.rb +21 -6
- data/lib/scout/llm/chat/process/meta.rb +421 -34
- data/lib/scout/llm/chat/process/options.rb +21 -1
- data/lib/scout/llm/chat/process/tools.rb +56 -15
- data/lib/scout/llm/chat/process.rb +4 -0
- data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
- data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
- data/lib/scout/llm/chat/prompt.rb +48 -0
- data/lib/scout/llm/chat/provenance.rb +775 -0
- data/lib/scout/llm/chat/tool_calls.rb +76 -0
- data/lib/scout/llm/chat.rb +18 -2
- data/lib/scout/llm/embed.rb +11 -3
- data/lib/scout/llm/image.rb +86 -0
- data/lib/scout/llm/mcp.rb +10 -2
- data/lib/scout/llm/rag.rb +3 -3
- data/lib/scout/llm/tools/call.rb +160 -11
- data/lib/scout/llm/tools/knowledge_base.rb +1 -1
- data/lib/scout/llm/tools/workflow.rb +32 -16
- data/lib/scout/model/python/huggingface/causal.rb +23 -5
- data/lib/scout/model/python/huggingface.rb +2 -1
- data/lib/scout-ai.rb +1 -0
- data/python/README.md +197 -14
- data/python/scout_ai/huggingface/eval.py +245 -34
- data/python/tests/test_huggingface_eval.py +58 -0
- data/research/ChatAnalyst-required-changes.md +167 -0
- data/research/agent-delegation-analysis.md +810 -0
- data/research/agent-meta-provenance-integration-plan.md +622 -0
- data/research/agent-workflow-analysis.md +1120 -0
- data/research/backends-analysis.md +836 -0
- data/research/chat-core-analysis.md +946 -0
- data/research/chatanalyst-provenance/00-baseline.md +30 -0
- data/research/chatanalyst-provenance/01-repo-map.md +60 -0
- data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
- data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
- data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
- data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
- data/research/chatanalyst-provenance/07-critic-review.md +25 -0
- data/research/chatanalyst-provenance/final-report.md +45 -0
- data/research/chatanalyst-provenance/resumption.md +37 -0
- data/research/coding-philosophy-analysis.md +928 -0
- data/research/commands-analysis.md +947 -0
- data/research/multi-agent-patterns-analysis.md +853 -0
- data/research/prompt-strategies-analysis.md +630 -0
- data/research/prov-verbosity-fix-notes.md +77 -0
- data/research/provenance-analysis.md +469 -0
- data/research/provenance-navigation-design.md +640 -0
- data/research/synthesis-report.md +487 -0
- data/research/tools-system-analysis.md +779 -0
- data/scout-ai.gemspec +100 -11
- data/scout_commands/agent/ask +13 -3
- data/scout_commands/agent/kb +2 -0
- data/scout_commands/llm/ask +11 -4
- data/scout_commands/llm/md +76 -0
- data/scout_commands/llm/process_queries +48 -0
- data/scout_commands/llm/prov +602 -0
- data/scout_commands/llm/word +71 -0
- data/scout_commands/workflow/mcp +43 -0
- data/share/word/reference.docx +0 -0
- data/test/etc/AI/mock.yaml +11 -0
- data/test/fixtures/backends/anthropic.json +19 -0
- data/test/fixtures/backends/anthropic_tool_use.json +24 -0
- data/test/fixtures/backends/bedrock.json +8 -0
- data/test/fixtures/backends/bedrock_embedding.json +3 -0
- data/test/fixtures/backends/bedrock_tool_use.json +17 -0
- data/test/fixtures/backends/ollama.json +16 -0
- data/test/fixtures/backends/ollama_tool_call.json +27 -0
- data/test/fixtures/backends/openai_chat.json +21 -0
- data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
- data/test/fixtures/backends/responses.json +33 -0
- data/test/fixtures/backends/responses_tool_call.json +28 -0
- data/test/integration/README.md +32 -0
- data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
- data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
- data/test/integration/scout/llm/backends/test_relay.rb +52 -0
- data/test/integration/scout/llm/test_infrastructure.rb +74 -0
- data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
- data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
- data/test/integration/scout/model/test_base.rb +91 -0
- data/test/scout/llm/agent/test_chat.rb +8 -2
- data/test/scout/llm/agent/test_save.rb +413 -0
- data/test/scout/llm/agent/test_workflow.rb +110 -0
- data/test/scout/llm/backends/test_anthropic.rb +93 -10
- data/test/scout/llm/backends/test_bedrock.rb +118 -2
- data/test/scout/llm/backends/test_huggingface.rb +137 -42
- data/test/scout/llm/backends/test_ollama.rb +70 -20
- data/test/scout/llm/backends/test_openwebui.rb +42 -40
- data/test/scout/llm/backends/test_relay.rb +4 -2
- data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
- data/test/scout/llm/chat/process/test_meta.rb +518 -0
- data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
- data/test/scout/llm/chat/test_agent_meta.rb +357 -0
- data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
- data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
- data/test/scout/llm/chat/test_parse.rb +70 -15
- data/test/scout/llm/chat/test_prov_cli.rb +274 -0
- data/test/scout/llm/chat/test_provenance.rb +240 -0
- data/test/scout/llm/chat/test_tool_calls.rb +38 -0
- data/test/scout/llm/test_agent.rb +13 -36
- data/test/scout/llm/test_ask.rb +75 -52
- data/test/scout/llm/test_chat.rb +107 -13
- data/test/scout/llm/test_embed.rb +48 -0
- data/test/scout/llm/test_rag.rb +23 -16
- data/test/scout/llm/test_tools.rb +12 -1
- data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
- data/test/scout/llm/tools/test_mcp.rb +5 -3
- data/test/scout/llm/tools/test_workflow.rb +23 -2
- data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
- data/test/scout/model/python/huggingface/test_causal.rb +9 -3
- data/test/scout/model/python/huggingface/test_classification.rb +11 -2
- data/test/scout/model/python/test_torch.rb +2 -0
- data/test/scout/model/python/torch/test_helpers.rb +4 -0
- data/test/scout/model/test_base.rb +4 -2
- data/test/support/availability.rb +231 -0
- data/test/support/fake_clients.rb +138 -0
- data/test/support/fixtures.rb +21 -0
- data/test/support/infrastructure_probes.rb +136 -0
- data/test/support/mock_backend.rb +215 -0
- data/test/test_helper.rb +32 -2
- metadata +99 -10
- data/doc/Agent.md +0 -327
- data/doc/Chat.md +0 -458
- data/doc/LLM.md +0 -340
- data/doc/RAG.md +0 -129
- data/scout_commands/documenter +0 -148
- data/test/scout/llm/backends/test_openai.rb +0 -192
- data/test/scout/llm/backends/test_responses.rb +0 -238
- data/test/scout/llm/test_parse.rb +0 -98
|
@@ -0,0 +1,640 @@
|
|
|
1
|
+
# Provenance navigation in Scout-AI
|
|
2
|
+
|
|
3
|
+
> This is an investigation and design proposal, not normative documentation.
|
|
4
|
+
> It reflects the code inspected in `scout-ai` and the current
|
|
5
|
+
> `/bulk/mvazque2/git/workflows/ChatAnalyst` checkout. It intentionally does
|
|
6
|
+
> not propose implementations for agent monitoring, `ask_user`, or explicit
|
|
7
|
+
> agent-raised failure; those topics are considered only where they constrain
|
|
8
|
+
> provenance.
|
|
9
|
+
|
|
10
|
+
## Executive conclusion
|
|
11
|
+
|
|
12
|
+
The difficult part of Scout-AI provenance is not recursion. It is that a run
|
|
13
|
+
interleaves two already-valid Scout data models:
|
|
14
|
+
|
|
15
|
+
1. a Chat is an Array of messages and contains conversational lineage,
|
|
16
|
+
inference metadata, imports, and job projections;
|
|
17
|
+
2. a Step is a persisted Workflow execution and contains dependencies, status,
|
|
18
|
+
inputs, result files, and log artifacts.
|
|
19
|
+
|
|
20
|
+
The provenance structure is therefore a heterogeneous directed graph whose
|
|
21
|
+
only persistent node kinds are **chat files** and **workflow jobs**. Agents,
|
|
22
|
+
tool calls, inference segments, and token events are observations inside those
|
|
23
|
+
nodes; they should not be promoted into Session, ProvenanceContext, ChatGraph,
|
|
24
|
+
or Node classes merely to make traversal convenient.
|
|
25
|
+
|
|
26
|
+
The core library should provide one small, block-oriented traversal primitive
|
|
27
|
+
that walks native `Path` and `Step` values and yields typed relationships. The
|
|
28
|
+
`prov` command and ChatAnalyst should consume that primitive and independently
|
|
29
|
+
format or analyse what they receive. Message tracing, tool-call pairing, and
|
|
30
|
+
token accounting should remain separate Chat operations layered on top of the
|
|
31
|
+
files returned by traversal.
|
|
32
|
+
|
|
33
|
+
The most important data-model gap is more fundamental: direct inference meta
|
|
34
|
+
messages currently have no persisted unique inference/event identifier. The
|
|
35
|
+
lineage digest is excellent for recognizing copied conversation history, but
|
|
36
|
+
it cannot distinguish two genuinely repeated requests with identical history
|
|
37
|
+
and metadata. Consequently, current token deduplication is useful but cannot
|
|
38
|
+
be exact in every case. A locally generated inference ID should eventually be
|
|
39
|
+
persisted in each direct meta record. This is independent of traversal and
|
|
40
|
+
should not block the traversal cleanup.
|
|
41
|
+
|
|
42
|
+
## Sources examined
|
|
43
|
+
|
|
44
|
+
The investigation covered:
|
|
45
|
+
|
|
46
|
+
- `scout_commands/llm/prov`
|
|
47
|
+
- `lib/scout/llm/chat/process/meta.rb`
|
|
48
|
+
- `lib/scout/llm/chat/provenance.rb`
|
|
49
|
+
- `lib/scout/llm/agent/workflow.rb`
|
|
50
|
+
- `lib/scout/llm/agent/delegate.rb`
|
|
51
|
+
- `lib/scout/llm/tools/call.rb`
|
|
52
|
+
- `lib/scout/llm/backends/default.rb`
|
|
53
|
+
- provenance and meta tests under `test/scout/llm/`
|
|
54
|
+
- the current and previous ChatAnalyst implementations in
|
|
55
|
+
`/bulk/mvazque2/git/workflows/ChatAnalyst/workflow.rb` and
|
|
56
|
+
`workflow.rb.orig`
|
|
57
|
+
- persisted jobs and recursively nested agent logs under
|
|
58
|
+
`~/.scout/var/jobs`
|
|
59
|
+
- agent examples under `~/chats/Agent`
|
|
60
|
+
- Scout Workflow lifecycle and provenance documentation
|
|
61
|
+
|
|
62
|
+
A real `Refined/ask` result with separate
|
|
63
|
+
`log/worker-round-1/agent.chat` and `log/critic-round-1/agent.chat` files was
|
|
64
|
+
used as a concrete layout check. The current ChatAnalyst Session found five
|
|
65
|
+
chats, two jobs, and six edges from one sampled result, which confirms that the
|
|
66
|
+
basic recursive idea works. It does not remove the design and correctness
|
|
67
|
+
issues below.
|
|
68
|
+
|
|
69
|
+
## The crisp model
|
|
70
|
+
|
|
71
|
+
### Persistent facts
|
|
72
|
+
|
|
73
|
+
A provenance reader should begin with facts that are already persisted, not
|
|
74
|
+
with inferred conceptual objects.
|
|
75
|
+
|
|
76
|
+
A chat file can directly state:
|
|
77
|
+
|
|
78
|
+
- conversational messages;
|
|
79
|
+
- `import`, `continue`, and `last` references to other chat files;
|
|
80
|
+
- direct inference metadata in `meta` messages;
|
|
81
|
+
- a producer job reference in `meta job=...`;
|
|
82
|
+
- function calls and function-call outputs.
|
|
83
|
+
|
|
84
|
+
A workflow job can directly state or expose:
|
|
85
|
+
|
|
86
|
+
- its path and logical Workflow/task identity;
|
|
87
|
+
- status, timing, exception, inputs, and dependency paths in Step info;
|
|
88
|
+
- its result, which may itself be a chat file;
|
|
89
|
+
- recursively named chat artifacts under `.files/log/**/*.chat`.
|
|
90
|
+
|
|
91
|
+
Nothing else is required to navigate the structural provenance.
|
|
92
|
+
|
|
93
|
+
### Structural relations
|
|
94
|
+
|
|
95
|
+
Traversal from a requested root follows five relations:
|
|
96
|
+
|
|
97
|
+
| Parent | Relation | Child | Meaning |
|
|
98
|
+
|---|---|---|---|
|
|
99
|
+
| chat | `import` | chat | The parent incorporates another persisted chat. |
|
|
100
|
+
| chat | `job` | job | A projected response in the chat was produced by the job. |
|
|
101
|
+
| job | `dependency` | job | The job consumes a normal Scout dependency. |
|
|
102
|
+
| job | `log` | chat | The job persisted an agent conversation as a log artifact. |
|
|
103
|
+
| job | `result` | chat | The job result itself is a chat and must be inspected, especially when traversal starts from a Step. |
|
|
104
|
+
|
|
105
|
+
These are **root-outward discovery relations**. They are deliberately not
|
|
106
|
+
presentation arrows. For example, a flow diagram may render a producer job as
|
|
107
|
+
feeding a chat, while recursive discovery follows the chat's `job` reference
|
|
108
|
+
toward the producer. `prov` currently mixes discovery hierarchy with natural
|
|
109
|
+
data-flow direction, which makes otherwise simple code hard to reason about.
|
|
110
|
+
Renderers should decide arrow direction after traversal.
|
|
111
|
+
|
|
112
|
+
### Observations inside nodes
|
|
113
|
+
|
|
114
|
+
The following are analyses of node content, not additional structural node
|
|
115
|
+
kinds:
|
|
116
|
+
|
|
117
|
+
- an inference segment is a meta record plus the messages it covers;
|
|
118
|
+
- a tool invocation is a paired `function_call` and
|
|
119
|
+
`function_call_output`;
|
|
120
|
+
- agent delegation is a tool invocation whose tool is `ask` or
|
|
121
|
+
`hand_off_to_*`;
|
|
122
|
+
- an agent log is still a chat file; “agent” is a label inferred from its log
|
|
123
|
+
path or explicit future metadata;
|
|
124
|
+
- token usage is attached to direct inference segments;
|
|
125
|
+
- a job failure or abort is Step status and exception evidence.
|
|
126
|
+
|
|
127
|
+
Keeping this distinction prevents the graph from becoming a second runtime
|
|
128
|
+
object model.
|
|
129
|
+
|
|
130
|
+
## What currently works well
|
|
131
|
+
|
|
132
|
+
Several core decisions are strong and should be retained.
|
|
133
|
+
|
|
134
|
+
### Chat remains plain data
|
|
135
|
+
|
|
136
|
+
`Chat.load` reads persisted chat text without recompiling roles. This is the
|
|
137
|
+
right safety boundary for provenance: inspection must never execute imports,
|
|
138
|
+
tasks, files, or tools.
|
|
139
|
+
|
|
140
|
+
### Job projections are explicit
|
|
141
|
+
|
|
142
|
+
`Chat.project` removes direct meta records from the projected output and adds a
|
|
143
|
+
single `meta job=<path>` marker. This cleanly distinguishes “tokens spent in
|
|
144
|
+
this chat” from “content produced elsewhere and projected here.”
|
|
145
|
+
|
|
146
|
+
### Conversational lineage is content-addressed
|
|
147
|
+
|
|
148
|
+
`message_index` hashes each non-meta message with its conversational
|
|
149
|
+
predecessor. Copied or inherited histories therefore share lineage identities,
|
|
150
|
+
while changed histories diverge. Excluding meta from the conversational chain
|
|
151
|
+
is correct because meta is not provider input.
|
|
152
|
+
|
|
153
|
+
### Segment tracing is independent of job traversal
|
|
154
|
+
|
|
155
|
+
`trace_indices`, `trace_chats`, `direct_entries`, and `token_totals` operate on
|
|
156
|
+
Chats and do not need to know how files were found. This separation is exactly
|
|
157
|
+
the direction the rest of the provenance implementation should follow.
|
|
158
|
+
|
|
159
|
+
### Workflow provenance is already authoritative for jobs
|
|
160
|
+
|
|
161
|
+
Step info already records dependencies, lifecycle status, timings, exceptions,
|
|
162
|
+
and other execution facts. Scout-AI should link to that evidence rather than
|
|
163
|
+
replicate it in a Session-like object.
|
|
164
|
+
|
|
165
|
+
## Current problems
|
|
166
|
+
|
|
167
|
+
### Traversal is duplicated three times
|
|
168
|
+
|
|
169
|
+
The same heterogeneous walk is independently implemented in:
|
|
170
|
+
|
|
171
|
+
- `scout_commands/llm/prov#report`;
|
|
172
|
+
- `Chat.provenance` and related helpers;
|
|
173
|
+
- `ChatAnalyst::Session` (and previously the much larger
|
|
174
|
+
`ChatAnalyst::ChatGraph`).
|
|
175
|
+
|
|
176
|
+
Each copy makes slightly different decisions about imports, result chats,
|
|
177
|
+
dependencies, logs, deduplication, errors, and edge direction. Fixes therefore
|
|
178
|
+
do not propagate to all consumers.
|
|
179
|
+
|
|
180
|
+
### Core helper contracts are inconsistent
|
|
181
|
+
|
|
182
|
+
Current examples include:
|
|
183
|
+
|
|
184
|
+
- `Chat.job_chat_files(job)` recursively includes dependencies;
|
|
185
|
+
- `Chat.job_agent_chat_files(job)` includes only the given job's logs;
|
|
186
|
+
- the instance `job_agent_chat_files` only expands the chat's immediate jobs;
|
|
187
|
+
- documentation describes some of these as recursively traversing all
|
|
188
|
+
dependencies;
|
|
189
|
+
- `job_agent_chats` is named as if it returned loaded Chat values but currently
|
|
190
|
+
returns paths.
|
|
191
|
+
|
|
192
|
+
A caller cannot infer recursion or return type from these names reliably.
|
|
193
|
+
Traversal should be centralized, and direct-neighbour helpers should say that
|
|
194
|
+
they are direct.
|
|
195
|
+
|
|
196
|
+
### `Chat.provenance` is not a faithful general traversal
|
|
197
|
+
|
|
198
|
+
The current implementation follows job logs and then calls
|
|
199
|
+
`Chat.provenance(dep.path)` for dependencies. That treats an arbitrary Step
|
|
200
|
+
result path as if it were necessarily a chat. It also omits chat imports and
|
|
201
|
+
uses a nested Hash whose meaning is limited to chat-to-log relationships;
|
|
202
|
+
normal dependency and producer relationships are not represented explicitly.
|
|
203
|
+
Broad existence checks and implicit rescues obscure malformed or missing
|
|
204
|
+
evidence.
|
|
205
|
+
|
|
206
|
+
### `prov` combines discovery, printing, graph construction, token analysis,
|
|
207
|
+
filtering, naming, and rendering
|
|
208
|
+
|
|
209
|
+
The recursive `report` both prints and builds a nested graph. Later flow code
|
|
210
|
+
must reverse or reinterpret edges based on whether a Hash key happens to be a
|
|
211
|
+
Step or a String. Global instance variables cache nodes, edges, chats, and
|
|
212
|
+
tokens. This makes the command longer and less reliable than the underlying
|
|
213
|
+
problem warrants.
|
|
214
|
+
|
|
215
|
+
There are also concrete fragilities:
|
|
216
|
+
|
|
217
|
+
- identity alternates between Step objects and paths;
|
|
218
|
+
- `report` may return `nil` for a seen object while callers expect a Hash to
|
|
219
|
+
merge;
|
|
220
|
+
- imports are absent;
|
|
221
|
+
- an unused `load_chat` helper remains;
|
|
222
|
+
- root detection relies on the presence of `<filename>.files`;
|
|
223
|
+
- display suppression of `agent.chat` is entangled with token attribution;
|
|
224
|
+
- “agent” type is guessed from `'.files/log/'` in a path.
|
|
225
|
+
|
|
226
|
+
### ChatAnalyst's Session is a cache plus recursive side effects
|
|
227
|
+
|
|
228
|
+
`Session` eagerly resolves and traverses everything in its constructor, stores
|
|
229
|
+
four mutable collections, embeds path resolution and warning policy, and adds
|
|
230
|
+
token methods that already exist on Chat. Every task constructs a fresh
|
|
231
|
+
Session and then reads those caches.
|
|
232
|
+
|
|
233
|
+
This abstraction adds no domain concept. It is an execution context for one
|
|
234
|
+
algorithm. More importantly, it encourages future agent-written code to add
|
|
235
|
+
more convenience methods until traversal, accounting, tool semantics, and
|
|
236
|
+
reporting are coupled again.
|
|
237
|
+
|
|
238
|
+
The earlier `ChatGraph` demonstrates that failure mode clearly: it grew path
|
|
239
|
+
canonicalization, hidden-path policy, tool parsing, success inference, usage
|
|
240
|
+
accounting, overlap analysis, message storage, job fallbacks, and graph
|
|
241
|
+
construction inside one class. Replacing it with a smaller Session reduced the
|
|
242
|
+
amount of code but not the architectural cause.
|
|
243
|
+
|
|
244
|
+
### Error handling is mostly invisible
|
|
245
|
+
|
|
246
|
+
Some library traversal helpers rescue all exceptions and return empty arrays.
|
|
247
|
+
ChatAnalyst catches exceptions and stores warning strings. `prov` often lets
|
|
248
|
+
errors escape. These policies make the same missing log or unreadable legacy
|
|
249
|
+
Step look like “no provenance,” a warning, or a fatal error depending on the
|
|
250
|
+
consumer.
|
|
251
|
+
|
|
252
|
+
The traversal primitive should not invent an error-monitoring subsystem, but
|
|
253
|
+
it must expose failures with their source node and attempted relation so that a
|
|
254
|
+
CLI can warn, an analyst can report incomplete evidence, and a strict caller
|
|
255
|
+
can raise.
|
|
256
|
+
|
|
257
|
+
### Location identity and lineage identity are conflated
|
|
258
|
+
|
|
259
|
+
There are two legitimate identities:
|
|
260
|
+
|
|
261
|
+
- a **message address**, such as `[chat_path, index]`, identifies where a
|
|
262
|
+
persisted message can be retrieved;
|
|
263
|
+
- a **lineage ID** identifies equivalent conversational content and is useful
|
|
264
|
+
for recognizing inherited or copied history.
|
|
265
|
+
|
|
266
|
+
ChatAnalyst creates strings such as `path#index` and calls them IDs, while Chat
|
|
267
|
+
also calls the lineage digest an ID. Reports should name these fields
|
|
268
|
+
`address` and `lineage_id` explicitly. An address should remain structured
|
|
269
|
+
until final JSON formatting rather than relying on parsing a path containing a
|
|
270
|
+
separator.
|
|
271
|
+
|
|
272
|
+
### Exact inference identity is not persisted
|
|
273
|
+
|
|
274
|
+
`Backend::Default#update_meta` currently stores token fields and cumulative
|
|
275
|
+
snapshots, but the tests explicitly assert that `usage_id` is absent.
|
|
276
|
+
`trace_chats` deduplicates segments by the lineage-derived meta ID. This is
|
|
277
|
+
correct for a copied historical segment but can collapse two genuinely
|
|
278
|
+
repeated requests when all of the following are identical:
|
|
279
|
+
|
|
280
|
+
- preceding conversational lineage;
|
|
281
|
+
- meta token fields;
|
|
282
|
+
- response content.
|
|
283
|
+
|
|
284
|
+
Conversely, session counters are process snapshots and cannot identify a
|
|
285
|
+
request. Exact cost and execution counting requires a unique persisted direct
|
|
286
|
+
inference identity generated by Scout-AI, independent of provider request IDs.
|
|
287
|
+
|
|
288
|
+
### Delegation is not always structurally linked to its log
|
|
289
|
+
|
|
290
|
+
Workflow-backed agent calls have strong job and log links. Socialized agents
|
|
291
|
+
are returned by the `ask` tool, executed later by `process_calls`, and finally
|
|
292
|
+
written by `log_agent` under a society path. The function call says which
|
|
293
|
+
agent/conversation was requested, and the path convention often allows a
|
|
294
|
+
match, but no explicit durable ID links that call to the precise specialist
|
|
295
|
+
chat or inference segment. This should be reported as a semantic association,
|
|
296
|
+
not presented as an authoritative structural edge.
|
|
297
|
+
|
|
298
|
+
## Proposed core primitives
|
|
299
|
+
|
|
300
|
+
### 1. Direct-neighbour readers
|
|
301
|
+
|
|
302
|
+
First make direct facts explicit and non-recursive. Suggested responsibilities
|
|
303
|
+
are:
|
|
304
|
+
|
|
305
|
+
- resolve a persisted chat reference without compiling it;
|
|
306
|
+
- return direct import files for a chat and its source path;
|
|
307
|
+
- return direct job references from a chat;
|
|
308
|
+
- return direct dependencies of a Step;
|
|
309
|
+
- return direct log chat files of a Step;
|
|
310
|
+
- return the Step result chat when its type is chat.
|
|
311
|
+
|
|
312
|
+
Existing APIs can supply several of these, but names and contracts should be
|
|
313
|
+
made consistent. In particular, methods named `*_chats` should return Chat
|
|
314
|
+
values and methods named `*_files` should return paths. Recursion should not be
|
|
315
|
+
hidden inside either.
|
|
316
|
+
|
|
317
|
+
### 2. One block-oriented traversal
|
|
318
|
+
|
|
319
|
+
Add a module function on Chat, not a class. A possible contract is:
|
|
320
|
+
|
|
321
|
+
Chat.traverse_provenance(root, follow: :all, on_error: nil) do |
|
|
322
|
+
kind, object, parent_kind, parent, relation
|
|
323
|
+
|
|
|
324
|
+
# kind is :chat or :job
|
|
325
|
+
# object is a Path for :chat and a Step for :job
|
|
326
|
+
# root has nil parent and relation
|
|
327
|
+
end
|
|
328
|
+
|
|
329
|
+
Without a block it should return an Enumerator. The implementation needs only
|
|
330
|
+
a queue or recursion and a Set. The visited key must include node kind, because
|
|
331
|
+
a chat-typed job result may have the same filesystem path as its Step.
|
|
332
|
+
|
|
333
|
+
The traversal should:
|
|
334
|
+
|
|
335
|
+
1. normalize a root explicitly as a chat file or Step;
|
|
336
|
+
2. yield the root once;
|
|
337
|
+
3. read only direct neighbours;
|
|
338
|
+
4. yield each newly visited child together with the relation that discovered
|
|
339
|
+
it;
|
|
340
|
+
5. retain cycle safety and shared dependency deduplication;
|
|
341
|
+
6. preserve native values rather than wrapping them in node classes;
|
|
342
|
+
7. expose read/load failures through `on_error` with the parent and relation.
|
|
343
|
+
|
|
344
|
+
The exact block argument order can change, but the essential contract is that
|
|
345
|
+
consumers receive native objects plus explicit type and relation. A flat stream
|
|
346
|
+
is sufficient to print a tree, collect JSON nodes and edges, calculate tokens,
|
|
347
|
+
or produce DOT.
|
|
348
|
+
|
|
349
|
+
An error callback can receive:
|
|
350
|
+
|
|
351
|
+
error, kind, object, relation, referenced_value
|
|
352
|
+
|
|
353
|
+
If no callback is supplied, raising is the clearest library default. CLI and
|
|
354
|
+
analysis callers can opt into “record warning and continue.” Silent rescue and
|
|
355
|
+
empty output should not be the default because absence of evidence differs
|
|
356
|
+
from unreadable evidence.
|
|
357
|
+
|
|
358
|
+
### 3. Small collectors implemented from traversal
|
|
359
|
+
|
|
360
|
+
Convenience is still useful when it preserves the same semantics. Small module
|
|
361
|
+
functions can collect from the stream:
|
|
362
|
+
|
|
363
|
+
- `Chat.provenance_chat_files(root)`;
|
|
364
|
+
- `Chat.provenance_jobs(root)`;
|
|
365
|
+
- `Chat.provenance_edges(root)`.
|
|
366
|
+
|
|
367
|
+
These should be thin collectors, not alternate traversal implementations.
|
|
368
|
+
The existing `Chat.provenance` can either become a compatibility collector or
|
|
369
|
+
be deprecated after consumers migrate.
|
|
370
|
+
|
|
371
|
+
### 4. Source-aware tracing
|
|
372
|
+
|
|
373
|
+
Keep `Chat.trace_chats` for compatibility, but add a source-aware form that
|
|
374
|
+
accepts `[path, chat]` pairs and retains addresses:
|
|
375
|
+
|
|
376
|
+
{ meta_address: [path, index],
|
|
377
|
+
lineage_id: ...,
|
|
378
|
+
meta: ...,
|
|
379
|
+
messages: [[path, index], ...],
|
|
380
|
+
message_lineages: [...] }
|
|
381
|
+
|
|
382
|
+
Deduplication should remain based on explicit inference ID when present and
|
|
383
|
+
fall back to the current lineage rule for legacy chats. Retrieval should use
|
|
384
|
+
addresses; overlap analysis should use lineage IDs.
|
|
385
|
+
|
|
386
|
+
### 5. Tool-call pairing as a Chat primitive
|
|
387
|
+
|
|
388
|
+
Pairing calls and outputs is message analysis used by any future analyst, not
|
|
389
|
+
only ChatAnalyst. Add a small operation that returns plain Hash records and
|
|
390
|
+
preserves both addresses. It should support current provider-normalized roles
|
|
391
|
+
and fields (`function_call`, `mcp_call`, `function_call_output`, `id`,
|
|
392
|
+
`call_id`, nested function name).
|
|
393
|
+
|
|
394
|
+
Success/failure interpretation should be separate. Pairing can authoritatively
|
|
395
|
+
say whether an output exists and what it contains. Deciding that JSON with
|
|
396
|
+
`exit_status != 0` means a failed shell call is a reporting policy, not generic
|
|
397
|
+
Chat structure. A helper may provide the common policy, but the raw pair must
|
|
398
|
+
remain available.
|
|
399
|
+
|
|
400
|
+
## Suggested traversal implementation shape
|
|
401
|
+
|
|
402
|
+
The implementation can be short and procedural:
|
|
403
|
+
|
|
404
|
+
1. a queue contains tuples of kind, native object, parent kind, parent, and
|
|
405
|
+
relation;
|
|
406
|
+
2. a Set contains `[kind, canonical_path]` keys;
|
|
407
|
+
3. visiting a chat loads it once, enqueues direct imports and producer jobs;
|
|
408
|
+
4. visiting a job enqueues direct dependencies, logs, and its result chat;
|
|
409
|
+
5. each loading operation is wrapped only at the boundary needed to call the
|
|
410
|
+
configured error handler.
|
|
411
|
+
|
|
412
|
+
There is no need for Session state, a Graph object, node subclasses, edge
|
|
413
|
+
classes, visitor classes, or a provenance database.
|
|
414
|
+
|
|
415
|
+
A consumer that needs a graph can use ordinary Hashes and Arrays:
|
|
416
|
+
|
|
417
|
+
nodes = {}
|
|
418
|
+
edges = []
|
|
419
|
+
|
|
420
|
+
Chat.traverse_provenance(root, on_error: record_warning) do |
|
|
421
|
+
kind, object, parent_kind, parent, relation
|
|
422
|
+
|
|
|
423
|
+
key = [kind, provenance_path(object)]
|
|
424
|
+
nodes[key] ||= object
|
|
425
|
+
edges << [[parent_kind, provenance_path(parent)], relation, key] if parent
|
|
426
|
+
end
|
|
427
|
+
|
|
428
|
+
That state belongs to the report being built, not to the core traversal.
|
|
429
|
+
|
|
430
|
+
## How `prov` should change
|
|
431
|
+
|
|
432
|
+
`prov` should become three clearly separated stages.
|
|
433
|
+
|
|
434
|
+
### Discovery
|
|
435
|
+
|
|
436
|
+
Call `Chat.traverse_provenance` once and collect the flat visit stream or edges.
|
|
437
|
+
Do not print during recursion. Do not infer edge types from Ruby classes after
|
|
438
|
+
the fact.
|
|
439
|
+
|
|
440
|
+
### Analysis
|
|
441
|
+
|
|
442
|
+
Load each discovered chat once. Use Chat operations for direct token entries,
|
|
443
|
+
totals, and any future tool-call summaries. Job token totals can be defined as
|
|
444
|
+
the totals of chat logs directly owned by that job; subtree totals should be
|
|
445
|
+
named explicitly because they include dependencies and can overlap in a DAG.
|
|
446
|
+
|
|
447
|
+
The current command sometimes presents a job total that recursively includes
|
|
448
|
+
its logs and descendants without making scope obvious. Reports should label
|
|
449
|
+
`direct` versus `subtree` totals.
|
|
450
|
+
|
|
451
|
+
### Rendering
|
|
452
|
+
|
|
453
|
+
Tree, compact flow, DOT, and plots should consume the same node/edge arrays.
|
|
454
|
+
Tree indentation is a spanning-tree presentation of a DAG; repeated nodes
|
|
455
|
+
should be shown as references rather than recursively expanded. Flow arrow
|
|
456
|
+
direction should be selected by relation in rendering only.
|
|
457
|
+
|
|
458
|
+
This removes the global `@flow_*` caches and most type/path guesses from the
|
|
459
|
+
command.
|
|
460
|
+
|
|
461
|
+
## How ChatAnalyst should change
|
|
462
|
+
|
|
463
|
+
ChatAnalyst does not need Session.
|
|
464
|
+
|
|
465
|
+
Each task can use one Workflow helper that returns the traversal stream or a
|
|
466
|
+
plain collected Hash for the duration of that task. A helper is appropriate
|
|
467
|
+
because it is local executable reuse, not a new domain abstraction. For
|
|
468
|
+
example:
|
|
469
|
+
|
|
470
|
+
helper :provenance_records do |file, warnings = []|
|
|
471
|
+
Chat.traverse_provenance(file, on_error: ->(...) { warnings << ... }).to_a
|
|
472
|
+
end
|
|
473
|
+
|
|
474
|
+
Tasks can then remain focused:
|
|
475
|
+
|
|
476
|
+
- `message_index`: iterate discovered chat files and emit address plus lineage;
|
|
477
|
+
- `message_content`: retrieve exact addresses;
|
|
478
|
+
- `chat_overview`: collect structural nodes and edges;
|
|
479
|
+
- `chat_tool_calls`: call the shared Chat pairing primitive;
|
|
480
|
+
- `chat_tokens`: call source-aware trace/token operations;
|
|
481
|
+
- `chat_agents`: filter paired tool calls, while marking inferred log matches as
|
|
482
|
+
inferred;
|
|
483
|
+
- `chat_report`: use Workflow dependencies on the smaller tasks if caching is
|
|
484
|
+
desirable, or combine their plain helper results.
|
|
485
|
+
|
|
486
|
+
`chat_report` should not rerun a hidden second traversal for each subreport
|
|
487
|
+
inside one task. Either collect once locally or make reports proper Workflow
|
|
488
|
+
dependencies. That decision is normal Workflow composition, not provenance
|
|
489
|
+
architecture.
|
|
490
|
+
|
|
491
|
+
## Identity and deduplication rules
|
|
492
|
+
|
|
493
|
+
### Chats
|
|
494
|
+
|
|
495
|
+
For one filesystem, use an expanded real path where possible. Preserve the
|
|
496
|
+
original reference as an alias for reporting. Do not silently merge different
|
|
497
|
+
files merely because their contents overlap.
|
|
498
|
+
|
|
499
|
+
### Jobs
|
|
500
|
+
|
|
501
|
+
Use the loaded Step path as the physical identity. A logical identity such as
|
|
502
|
+
workflow, task, and result basename can identify mirrored `~/.scout` and
|
|
503
|
+
`~/.rbbt` candidates, but merging them is safe only if their relevant info and
|
|
504
|
+
result agree. If two physical jobs share a logical identity but disagree,
|
|
505
|
+
retain both and issue a warning rather than selecting one silently.
|
|
506
|
+
|
|
507
|
+
### Messages
|
|
508
|
+
|
|
509
|
+
Use `[chat_path, index]` as the address and the existing digest as
|
|
510
|
+
`lineage_id`.
|
|
511
|
+
|
|
512
|
+
### Inference events
|
|
513
|
+
|
|
514
|
+
Introduce a random locally generated `inference_id` (or equivalently named
|
|
515
|
+
field) in each direct meta record. Generate it once per actual backend request
|
|
516
|
+
and preserve it through cached replay and copied chat history. Provider request
|
|
517
|
+
IDs can be stored separately when available, but should not be required.
|
|
518
|
+
Legacy records fall back to lineage identity with an explicit
|
|
519
|
+
`deduplication: legacy_lineage` qualification in precise reports.
|
|
520
|
+
|
|
521
|
+
### Tool calls
|
|
522
|
+
|
|
523
|
+
Use provider/tool call ID scoped by chat lineage or chat address. Do not assume
|
|
524
|
+
a provider call ID is globally unique across all files.
|
|
525
|
+
|
|
526
|
+
## Correctness tests to add before migration
|
|
527
|
+
|
|
528
|
+
Traversal tests should use temporary persisted jobs/chats and cover:
|
|
529
|
+
|
|
530
|
+
1. a root chat with no references;
|
|
531
|
+
2. a chat importing another chat;
|
|
532
|
+
3. a chat projected from a job;
|
|
533
|
+
4. a job with multiple recursive log chats;
|
|
534
|
+
5. a job with a chat dependency and a non-chat dependency;
|
|
535
|
+
6. shared dependencies visited once but represented by all relevant edges;
|
|
536
|
+
7. a cycle caused by a chat result projecting its own producer job;
|
|
537
|
+
8. missing import, missing job, malformed chat, and unreadable Step info under
|
|
538
|
+
both strict and warning policies;
|
|
539
|
+
9. two physical job roots with the same logical digest;
|
|
540
|
+
10. an aborted or error Step with partial log files;
|
|
541
|
+
11. traversal starting from a Step rather than a chat;
|
|
542
|
+
12. deterministic traversal order.
|
|
543
|
+
|
|
544
|
+
Tracing/accounting tests should cover:
|
|
545
|
+
|
|
546
|
+
1. copied direct inference history deduplicated once;
|
|
547
|
+
2. two identical real requests with distinct `inference_id` counted twice;
|
|
548
|
+
3. legacy identical records reported with fallback deduplication;
|
|
549
|
+
4. job projection meta excluded from direct token totals;
|
|
550
|
+
5. cache, cache-write, and reasoning token fields;
|
|
551
|
+
6. source addresses retained for every segment and covered message;
|
|
552
|
+
7. orphan and consecutive meta records;
|
|
553
|
+
8. the same tool call ID appearing in two chat files;
|
|
554
|
+
9. missing tool output versus an output containing an exception;
|
|
555
|
+
10. socialized agent calls where a log association is only inferred.
|
|
556
|
+
|
|
557
|
+
Consumer tests should run `prov` and ChatAnalyst against the same fixture and
|
|
558
|
+
assert that they discover the same chat paths, jobs, and structural edges.
|
|
559
|
+
|
|
560
|
+
## Concrete defects worth fixing opportunistically
|
|
561
|
+
|
|
562
|
+
These are small enough to address while introducing the primitive:
|
|
563
|
+
|
|
564
|
+
- make `job_agent_chats` actually load and return Chat values, or rename it;
|
|
565
|
+
- document and test whether each `job_*` helper is direct or recursive;
|
|
566
|
+
- remove broad rescue-to-empty behavior from structural readers;
|
|
567
|
+
- remove the unused `load_chat` in `prov`;
|
|
568
|
+
- stop detecting job roots only through `<filename>.files`;
|
|
569
|
+
- avoid `merge!` on a recursive result that may be `nil` for a seen node;
|
|
570
|
+
- require `set` explicitly in `prov` if it continues to use Set directly;
|
|
571
|
+
- stop classifying an agent as a distinct persistent node based solely on a
|
|
572
|
+
path substring;
|
|
573
|
+
- make token scope explicit (`direct file`, `direct job logs`, or `subtree`);
|
|
574
|
+
- correct maintained documentation that currently claims recursive behavior or
|
|
575
|
+
APIs that do not match the code.
|
|
576
|
+
|
|
577
|
+
## Relationship to the deferred concerns
|
|
578
|
+
|
|
579
|
+
### Exceptions, aborted managers, and active-agent monitoring
|
|
580
|
+
|
|
581
|
+
The proposed traversal already helps historical failures because it must visit
|
|
582
|
+
jobs regardless of `done?`, expose Step status, and discover partial log files.
|
|
583
|
+
It should not require a new monitoring abstraction.
|
|
584
|
+
|
|
585
|
+
If active agents later write state under `var`, the clean integration is to
|
|
586
|
+
make that state another persisted artifact referenced by a Step or chat, not to
|
|
587
|
+
put live-agent state into the provenance walker. Historical provenance and
|
|
588
|
+
live monitoring have different consistency requirements.
|
|
589
|
+
|
|
590
|
+
Backend failure snapshots currently written to anonymous TmpFile paths are
|
|
591
|
+
weak provenance because the path is logged but not structurally linked to the
|
|
592
|
+
Step/chat. A future change could save the snapshot under the owning job's
|
|
593
|
+
files directory when `agent.job` is available, or record its path in Step info.
|
|
594
|
+
That would make it discoverable without changing traversal semantics. This is
|
|
595
|
+
not required for the provenance refactor.
|
|
596
|
+
|
|
597
|
+
### `ask_user`
|
|
598
|
+
|
|
599
|
+
An `ask_user` operation should appear as an ordinary tool call. Its spool item
|
|
600
|
+
and eventual response can carry the tool call ID or a derived request ID. The
|
|
601
|
+
generic tool-call pairing and provenance addresses proposed here are enough;
|
|
602
|
+
the walker should not special-case human interaction.
|
|
603
|
+
|
|
604
|
+
### Agents declaring impossible work
|
|
605
|
+
|
|
606
|
+
An explicit failure should become either a tool output containing structured
|
|
607
|
+
failure or a normal Workflow/Step exception and status. Again, traversal only
|
|
608
|
+
needs to preserve and expose that evidence. It should not decide whether the
|
|
609
|
+
failure was justified.
|
|
610
|
+
|
|
611
|
+
## Recommended implementation sequence
|
|
612
|
+
|
|
613
|
+
1. **Specify and test direct relations.** Correct helper names/return types and
|
|
614
|
+
add fixtures for chats, jobs, dependencies, logs, and imports.
|
|
615
|
+
2. **Add `Chat.traverse_provenance`.** Keep it procedural, block-oriented, and
|
|
616
|
+
explicit about errors.
|
|
617
|
+
3. **Reimplement existing collectors from traversal.** Preserve compatibility
|
|
618
|
+
where inexpensive.
|
|
619
|
+
4. **Migrate `prov`.** Separate discovery, analysis, and rendering; compare
|
|
620
|
+
output against current real jobs.
|
|
621
|
+
5. **Add source-aware tracing and tool-call pairing.** Keep plain Hash/Array
|
|
622
|
+
results.
|
|
623
|
+
6. **Migrate ChatAnalyst and delete Session.** Use Workflow helpers and tasks,
|
|
624
|
+
not replacement classes.
|
|
625
|
+
7. **Add persisted inference IDs.** Update accounting to prefer them while
|
|
626
|
+
retaining legacy lineage fallback.
|
|
627
|
+
8. **Revise maintained developer documentation.** Keep this investigation as
|
|
628
|
+
the detailed rationale and document only the stable concepts/API in
|
|
629
|
+
`doc/developer/Provenance.md`.
|
|
630
|
+
|
|
631
|
+
## Final design rule
|
|
632
|
+
|
|
633
|
+
A useful test for every proposed provenance abstraction is:
|
|
634
|
+
|
|
635
|
+
> Does this represent a persisted fact in Chat or Workflow, or is it only state
|
|
636
|
+
> needed while producing one report?
|
|
637
|
+
|
|
638
|
+
Persisted facts belong in Chat/Step primitives. Temporary report state belongs
|
|
639
|
+
in local Arrays, Hashes, Sets, and blocks. If an object exists only to hold a
|
|
640
|
+
queue, a seen set, loaded chats, edges, and warnings, it should not be a class.
|