scout-ai 1.2.3 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.vimproject +138 -50
- data/README.md +171 -290
- data/Rakefile +17 -1
- data/VERSION +1 -1
- data/doc/Improvements.md +325 -0
- data/doc/StartHere.md +110 -0
- data/doc/developer/Architecture.md +126 -0
- data/doc/developer/Backends.md +199 -0
- data/doc/developer/ChatLifecycle.md +183 -0
- data/doc/developer/DelegationInternals.md +295 -0
- data/doc/developer/DesignPrinciples.md +245 -0
- data/doc/developer/PromptProcessing.md +292 -0
- data/doc/developer/Provenance.md +317 -0
- data/doc/user/BuildingAgents.md +345 -0
- data/doc/user/Cookbook.md +333 -0
- data/doc/user/CoreConcepts.md +181 -0
- data/doc/user/Delegation.md +191 -0
- data/doc/user/GettingStarted.md +159 -0
- data/doc/user/ManagingContext.md +163 -0
- data/doc/user/MultiAgentWorkflows.md +256 -0
- data/doc/user/Python.md +159 -0
- data/doc/user/RunningInference.md +200 -0
- data/doc/user/ToolCalling.md +193 -0
- data/doc/user/WritingChats.md +197 -0
- data/lib/scout/llm/agent/chat.rb +61 -11
- data/lib/scout/llm/agent/delegate.rb +274 -65
- data/lib/scout/llm/agent/iterate.rb +2 -2
- data/lib/scout/llm/agent/save.rb +273 -0
- data/lib/scout/llm/agent/workflow.rb +164 -0
- data/lib/scout/llm/agent.rb +86 -61
- data/lib/scout/llm/ask.rb +62 -17
- data/lib/scout/llm/backends/anthropic.rb +9 -2
- data/lib/scout/llm/backends/bedrock.rb +15 -3
- data/lib/scout/llm/backends/default.rb +183 -99
- data/lib/scout/llm/backends/glm.rb +58 -0
- data/lib/scout/llm/backends/huggingface.rb +196 -26
- data/lib/scout/llm/backends/ollama.rb +13 -1
- data/lib/scout/llm/backends/openai.rb +0 -2
- data/lib/scout/llm/backends/openwebui.rb +20 -13
- data/lib/scout/llm/backends/relay.rb +22 -22
- data/lib/scout/llm/backends/responses.rb +1 -1
- data/lib/scout/llm/chat/agent_meta.rb +264 -0
- data/lib/scout/llm/chat/annotation.rb +39 -10
- data/lib/scout/llm/chat/parse.rb +28 -6
- data/lib/scout/llm/chat/persist.rb +25 -0
- data/lib/scout/llm/chat/process/clear.rb +41 -6
- data/lib/scout/llm/chat/process/files.rb +21 -6
- data/lib/scout/llm/chat/process/meta.rb +421 -34
- data/lib/scout/llm/chat/process/options.rb +21 -1
- data/lib/scout/llm/chat/process/tools.rb +56 -15
- data/lib/scout/llm/chat/process.rb +4 -0
- data/lib/scout/llm/chat/prompt/shorten_tools.rb +125 -0
- data/lib/scout/llm/chat/prompt/shorten_tools_epoch.rb +365 -0
- data/lib/scout/llm/chat/prompt.rb +48 -0
- data/lib/scout/llm/chat/provenance.rb +775 -0
- data/lib/scout/llm/chat/tool_calls.rb +76 -0
- data/lib/scout/llm/chat.rb +18 -2
- data/lib/scout/llm/embed.rb +11 -3
- data/lib/scout/llm/image.rb +86 -0
- data/lib/scout/llm/mcp.rb +10 -2
- data/lib/scout/llm/rag.rb +3 -3
- data/lib/scout/llm/tools/call.rb +160 -11
- data/lib/scout/llm/tools/knowledge_base.rb +1 -1
- data/lib/scout/llm/tools/workflow.rb +32 -16
- data/lib/scout/model/python/huggingface/causal.rb +23 -5
- data/lib/scout/model/python/huggingface.rb +2 -1
- data/lib/scout-ai.rb +1 -0
- data/python/README.md +197 -14
- data/python/scout_ai/huggingface/eval.py +245 -34
- data/python/tests/test_huggingface_eval.py +58 -0
- data/research/ChatAnalyst-required-changes.md +167 -0
- data/research/agent-delegation-analysis.md +810 -0
- data/research/agent-meta-provenance-integration-plan.md +622 -0
- data/research/agent-workflow-analysis.md +1120 -0
- data/research/backends-analysis.md +836 -0
- data/research/chat-core-analysis.md +946 -0
- data/research/chatanalyst-provenance/00-baseline.md +30 -0
- data/research/chatanalyst-provenance/01-repo-map.md +60 -0
- data/research/chatanalyst-provenance/02-event-reconstruction.md +55 -0
- data/research/chatanalyst-provenance/03-duplication-evidence.md +45 -0
- data/research/chatanalyst-provenance/04-tooling-root-cause.md +57 -0
- data/research/chatanalyst-provenance/05-fix-plan.md +46 -0
- data/research/chatanalyst-provenance/07-critic-review.md +25 -0
- data/research/chatanalyst-provenance/final-report.md +45 -0
- data/research/chatanalyst-provenance/resumption.md +37 -0
- data/research/coding-philosophy-analysis.md +928 -0
- data/research/commands-analysis.md +947 -0
- data/research/multi-agent-patterns-analysis.md +853 -0
- data/research/prompt-strategies-analysis.md +630 -0
- data/research/prov-verbosity-fix-notes.md +77 -0
- data/research/provenance-analysis.md +469 -0
- data/research/provenance-navigation-design.md +640 -0
- data/research/synthesis-report.md +487 -0
- data/research/tools-system-analysis.md +779 -0
- data/scout-ai.gemspec +100 -11
- data/scout_commands/agent/ask +13 -3
- data/scout_commands/agent/kb +2 -0
- data/scout_commands/llm/ask +11 -4
- data/scout_commands/llm/md +76 -0
- data/scout_commands/llm/process_queries +48 -0
- data/scout_commands/llm/prov +602 -0
- data/scout_commands/llm/word +71 -0
- data/scout_commands/workflow/mcp +43 -0
- data/share/word/reference.docx +0 -0
- data/test/etc/AI/mock.yaml +11 -0
- data/test/fixtures/backends/anthropic.json +19 -0
- data/test/fixtures/backends/anthropic_tool_use.json +24 -0
- data/test/fixtures/backends/bedrock.json +8 -0
- data/test/fixtures/backends/bedrock_embedding.json +3 -0
- data/test/fixtures/backends/bedrock_tool_use.json +17 -0
- data/test/fixtures/backends/ollama.json +16 -0
- data/test/fixtures/backends/ollama_tool_call.json +27 -0
- data/test/fixtures/backends/openai_chat.json +21 -0
- data/test/fixtures/backends/openai_chat_tool_call.json +31 -0
- data/test/fixtures/backends/responses.json +33 -0
- data/test/fixtures/backends/responses_tool_call.json +28 -0
- data/test/integration/README.md +32 -0
- data/test/integration/scout/llm/backends/test_endpoints.rb +34 -0
- data/test/integration/scout/llm/backends/test_openwebui.rb +61 -0
- data/test/integration/scout/llm/backends/test_relay.rb +52 -0
- data/test/integration/scout/llm/test_infrastructure.rb +74 -0
- data/test/{scout → integration/scout}/llm/test_mcp.rb +1 -1
- data/test/integration/scout/llm/tools/test_mcp.rb +42 -0
- data/test/integration/scout/model/test_base.rb +91 -0
- data/test/scout/llm/agent/test_chat.rb +8 -2
- data/test/scout/llm/agent/test_save.rb +413 -0
- data/test/scout/llm/agent/test_workflow.rb +110 -0
- data/test/scout/llm/backends/test_anthropic.rb +93 -10
- data/test/scout/llm/backends/test_bedrock.rb +118 -2
- data/test/scout/llm/backends/test_huggingface.rb +137 -42
- data/test/scout/llm/backends/test_ollama.rb +70 -20
- data/test/scout/llm/backends/test_openwebui.rb +42 -40
- data/test/scout/llm/backends/test_relay.rb +4 -2
- data/test/scout/llm/chat/agent_meta_fixtures.rb +131 -0
- data/test/scout/llm/chat/process/test_meta.rb +518 -0
- data/test/scout/llm/chat/process/test_normalize_usage.rb +183 -0
- data/test/scout/llm/chat/test_agent_meta.rb +357 -0
- data/test/scout/llm/chat/test_agent_meta_provenance.rb +467 -0
- data/test/scout/llm/chat/test_agent_meta_tokens.rb +594 -0
- data/test/scout/llm/chat/test_parse.rb +70 -15
- data/test/scout/llm/chat/test_prov_cli.rb +274 -0
- data/test/scout/llm/chat/test_provenance.rb +240 -0
- data/test/scout/llm/chat/test_tool_calls.rb +38 -0
- data/test/scout/llm/test_agent.rb +13 -36
- data/test/scout/llm/test_ask.rb +75 -52
- data/test/scout/llm/test_chat.rb +107 -13
- data/test/scout/llm/test_embed.rb +48 -0
- data/test/scout/llm/test_rag.rb +23 -16
- data/test/scout/llm/test_tools.rb +12 -1
- data/test/scout/llm/tools/test_knowledge_base.rb +0 -1
- data/test/scout/llm/tools/test_mcp.rb +5 -3
- data/test/scout/llm/tools/test_workflow.rb +23 -2
- data/test/scout/model/python/huggingface/causal/test_next_token.rb +11 -5
- data/test/scout/model/python/huggingface/test_causal.rb +9 -3
- data/test/scout/model/python/huggingface/test_classification.rb +11 -2
- data/test/scout/model/python/test_torch.rb +2 -0
- data/test/scout/model/python/torch/test_helpers.rb +4 -0
- data/test/scout/model/test_base.rb +4 -2
- data/test/support/availability.rb +231 -0
- data/test/support/fake_clients.rb +138 -0
- data/test/support/fixtures.rb +21 -0
- data/test/support/infrastructure_probes.rb +136 -0
- data/test/support/mock_backend.rb +215 -0
- data/test/test_helper.rb +32 -2
- metadata +99 -10
- data/doc/Agent.md +0 -327
- data/doc/Chat.md +0 -458
- data/doc/LLM.md +0 -340
- data/doc/RAG.md +0 -129
- data/scout_commands/documenter +0 -148
- data/test/scout/llm/backends/test_openai.rb +0 -192
- data/test/scout/llm/backends/test_responses.rb +0 -238
- data/test/scout/llm/test_parse.rb +0 -98
|
@@ -0,0 +1,77 @@
|
|
|
1
|
+
# prov verbosity fix notes
|
|
2
|
+
|
|
3
|
+
## Problem
|
|
4
|
+
|
|
5
|
+
The `scout-ai llm prov` command output was too verbose and token counts were
|
|
6
|
+
wrong.
|
|
7
|
+
|
|
8
|
+
Two issues:
|
|
9
|
+
|
|
10
|
+
1. **Verbosity**: The output showed many redundant nodes — `agent.chat` log
|
|
11
|
+
files, zero-token result chats, `(seen)` repeated nodes, inline relation
|
|
12
|
+
labels, and long filesystem paths.
|
|
13
|
+
|
|
14
|
+
2. **Token counts**: The root chat showed `total=0` because it had no direct
|
|
15
|
+
inference tokens. Every parent node should show the aggregate cost of its
|
|
16
|
+
entire provenance subtree.
|
|
17
|
+
|
|
18
|
+
## Root causes
|
|
19
|
+
|
|
20
|
+
### Job resolution
|
|
21
|
+
|
|
22
|
+
`Step.load` with a relative workflow reference like
|
|
23
|
+
`Planned/ask/Default_xyz.chat` resolved via `Path.find` to
|
|
24
|
+
`~/.scout/Planned/ask/Default_xyz.chat`, but the actual job data lives at
|
|
25
|
+
`~/.rbbt/var/jobs/Planned/ask/Default_xyz.chat`. The `.scout` directory did not
|
|
26
|
+
contain the job files, so traversal stopped immediately.
|
|
27
|
+
|
|
28
|
+
### Redundant nodes
|
|
29
|
+
|
|
30
|
+
The traversal visits chat files connected to jobs via four relations: `job`,
|
|
31
|
+
`dependency`, `log`, `result`. Of these:
|
|
32
|
+
|
|
33
|
+
- **`log` relation with `agent.chat`**: This is the chat produced by the
|
|
34
|
+
inference backend, stored at `<job>.files/log/agent.chat`. Its tokens are the
|
|
35
|
+
same tokens counted as the job's direct cost. Showing it as a separate node
|
|
36
|
+
duplicates the parent's cost.
|
|
37
|
+
- **`result` relation**: A chat-typed job's persisted result file IS the same
|
|
38
|
+
path as the job itself. The `[:chat, path]` node is a duplicate of the
|
|
39
|
+
`[:job, path]` node.
|
|
40
|
+
- **Repeated nodes in shared DAG branches**: When a dependency appears in
|
|
41
|
+
multiple branches, it was shown as `(seen)` in the old verbose output.
|
|
42
|
+
|
|
43
|
+
### Token aggregation
|
|
44
|
+
|
|
45
|
+
The old code only showed each node's **direct** token cost (from its own
|
|
46
|
+
`agent.chat` or log files). This meant a root chat with no direct inference
|
|
47
|
+
showed `total=0`. The correct behavior is to aggregate the entire subtree's
|
|
48
|
+
inference cost.
|
|
49
|
+
|
|
50
|
+
## Fixes applied
|
|
51
|
+
|
|
52
|
+
### `lib/scout/llm/chat/provenance.rb`
|
|
53
|
+
|
|
54
|
+
Added `Chat.load_job_reference(reference)` helper that tries the standard
|
|
55
|
+
`~/.rbbt/var/jobs/` storage location as a fallback when `Step.load` resolves to
|
|
56
|
+
a non-existent path. The traversal uses this for all `chat.jobs` references.
|
|
57
|
+
|
|
58
|
+
### `scout_commands/llm/prov`
|
|
59
|
+
|
|
60
|
+
1. **Hidden nodes**: `agent.chat` log files and result-relation chats are
|
|
61
|
+
skipped entirely in tree mode (not printed, not traversed into).
|
|
62
|
+
|
|
63
|
+
2. **No `(seen)` lines**: Repeated nodes are silently skipped in tree mode.
|
|
64
|
+
|
|
65
|
+
3. **No inline relation labels**: Tree lines show just `kind tokens label`.
|
|
66
|
+
|
|
67
|
+
4. **Shortened paths**: Log chats under jobs show their path relative to the
|
|
68
|
+
job's `.files/log/` directory (e.g. `society/Worker/cli_investigation.chat`).
|
|
69
|
+
|
|
70
|
+
5. **Aggregate mode** (default): Each node shows the total cost of all
|
|
71
|
+
reachable chat files in its provenance subtree, deduplicated by path.
|
|
72
|
+
|
|
73
|
+
6. **Component mode** (`--component`): Each node shows only its own direct
|
|
74
|
+
token cost.
|
|
75
|
+
|
|
76
|
+
7. **Flow and DOT modes**: Same hidden-node filtering applied. Flow mode shows
|
|
77
|
+
data-flow arrows (result/dependency/log) between visible nodes.
|
|
@@ -0,0 +1,469 @@
|
|
|
1
|
+
> **Disclaimer:** This is an architectural investigation, not normative
|
|
2
|
+
> documentation. It was produced during a documentation-revamp effort and may
|
|
3
|
+
> be outdated relative to the current codebase. Treat it as supporting
|
|
4
|
+
> reference material. For maintained documentation, see
|
|
5
|
+
> [../../doc/](../../doc/).
|
|
6
|
+
>
|
|
7
|
+
|
|
8
|
+
|
|
9
|
+
# Provenance System: Chat.provenance, trace_chats, prov/info Commands, and ChatAnalyst
|
|
10
|
+
|
|
11
|
+
> **Scope:** How Scout-AI records, traverses, and reports the lineage of every
|
|
12
|
+
> inference — across imported chats, ask-job results, nested agent logs, and
|
|
13
|
+
> dependency chains — and the two CLI commands (`prov` and `info`) and the
|
|
14
|
+
> ChatAnalyst workflow that consume that provenance data.
|
|
15
|
+
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
## Provenance Data Model
|
|
19
|
+
|
|
20
|
+
### What provenance metadata is stored
|
|
21
|
+
|
|
22
|
+
Scout-AI does not have a standalone "provenance" table. Instead, provenance
|
|
23
|
+
information is embedded in **meta messages** that appear inline within the
|
|
24
|
+
chat transcript itself. A meta message has `role: meta` and a content string
|
|
25
|
+
serialized as `key=value key=value ...`.
|
|
26
|
+
|
|
27
|
+
The serialization/parsing layer lives in
|
|
28
|
+
`lib/scout/llm/chat/process/meta.rb`:
|
|
29
|
+
|
|
30
|
+
| Method | Purpose |
|
|
31
|
+
|---|---|
|
|
32
|
+
| `Chat.serialize_meta(hash)` | Converts a hash to the `key=value` string format, sorted by value length. |
|
|
33
|
+
| `Chat.parse_meta(str)` | Parses a `key=value` string back into an `IndiferentHash`. |
|
|
34
|
+
| `Chat.meta(messages)` | Strips all `meta` messages from an array and returns a merged metadata hash. The last checkpoint with `pt_c`/`ct_c`/`tt_c` fields is used for cumulative counters. |
|
|
35
|
+
|
|
36
|
+
#### Fields that can appear in a meta message
|
|
37
|
+
|
|
38
|
+
**Token fields** (written by `Backend::Default#update_meta` in
|
|
39
|
+
`lib/scout/llm/backends/default.rb`):
|
|
40
|
+
|
|
41
|
+
| Field | Meaning |
|
|
42
|
+
|---|---|
|
|
43
|
+
| `pt` | Prompt tokens for this single inference |
|
|
44
|
+
| `ct` | Completion tokens for this single inference |
|
|
45
|
+
| `tt` | Total tokens for this single inference |
|
|
46
|
+
| `pt_s`, `ct_s`, `tt_s` | **Session** cumulative counters (per-thread running totals) |
|
|
47
|
+
| `pt_c`, `ct_c`, `tt_c` | **Chat** cumulative counters (persisted across requests) |
|
|
48
|
+
| `reas` | Reasoning summary string (truncated for display) |
|
|
49
|
+
|
|
50
|
+
**Job-reference field** (written when a chat-task result is projected):
|
|
51
|
+
|
|
52
|
+
| Field | Meaning |
|
|
53
|
+
|---|---|
|
|
54
|
+
| `job` | The canonical path of the Scout workflow job that produced this segment |
|
|
55
|
+
|
|
56
|
+
#### Two kinds of meta messages
|
|
57
|
+
|
|
58
|
+
1. **Direct inference meta** — contains `pt`, `ct`, `tt` (and optionally
|
|
59
|
+
cumulative/session variants and `reas`). This records one actual model call.
|
|
60
|
+
|
|
61
|
+
2. **Job projection meta** — contains only `job=<path>`. This marks a response
|
|
62
|
+
segment that was projected from an ask-workflow job. It has **zero direct
|
|
63
|
+
token cost** — the actual tokens are recorded in the job's own agent logs
|
|
64
|
+
and dependency chain.
|
|
65
|
+
|
|
66
|
+
### How annotation.rb tracks provenance
|
|
67
|
+
|
|
68
|
+
The file `lib/scout/llm/chat/annotation.rb` provides the `Chat` annotation
|
|
69
|
+
(through `extend Annotation`). It adds convenience methods like `user`,
|
|
70
|
+
`assistant`, `import`, `option`, `endpoint`, `model`, etc. But it does **not**
|
|
71
|
+
contain provenance-specific methods — those all live in
|
|
72
|
+
`lib/scout/llm/chat/process/meta.rb`.
|
|
73
|
+
|
|
74
|
+
Relevant provenance-related instance methods on `Chat` objects (all defined in
|
|
75
|
+
`process/meta.rb`):
|
|
76
|
+
|
|
77
|
+
| Method | Returns |
|
|
78
|
+
|---|---|
|
|
79
|
+
| `chat.job_paths` / `chat.jobs` | Array of `Path` objects extracted from all `meta job=...` messages |
|
|
80
|
+
| `chat.job_chat_files` | All chat files (result + logs) reachable from this chat's jobs and their dependencies |
|
|
81
|
+
| `chat.job_agent_chat_files` | All `log/**/*.chat` files from this chat's jobs and their dependencies |
|
|
82
|
+
| `chat.job_chats` | All `Chat` objects loaded from `job_chat_files` |
|
|
83
|
+
| `chat.job_agent_chats` | All `Chat` objects loaded from `job_agent_chat_files` |
|
|
84
|
+
| `chat.message_index` | Array of per-message lineage records with `id`, `role`, `prev`, `fingerprint`, and parsed `meta` |
|
|
85
|
+
| `chat.meta` | The parsed metadata hash from the last meta message |
|
|
86
|
+
| `chat.last_job` | The `job` value from the last meta message |
|
|
87
|
+
|
|
88
|
+
### The `Chat.project` method
|
|
89
|
+
|
|
90
|
+
When a chat-task produces output, `Chat.project(job, messages)` wraps the
|
|
91
|
+
non-meta messages with a single `meta job=<path>` marker at the front. This
|
|
92
|
+
ensures that consumers can detect the job origin without scanning for token
|
|
93
|
+
fields:
|
|
94
|
+
|
|
95
|
+
```ruby
|
|
96
|
+
[{ role: :meta, content: serialize_meta(job: job.to_s) }] + projected_messages
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
---
|
|
100
|
+
|
|
101
|
+
## Recursive Traversal / trace_chats
|
|
102
|
+
|
|
103
|
+
### Lineage IDs and message_index
|
|
104
|
+
|
|
105
|
+
`Chat#message_index` (defined in `process/meta.rb`) computes a **lineage ID**
|
|
106
|
+
for each message:
|
|
107
|
+
|
|
108
|
+
```ruby
|
|
109
|
+
id = Misc.digest([previous_id, role, content])
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
Each message's lineage ID incorporates the previous conversational message's ID
|
|
113
|
+
(meta messages are excluded from the lineage chain — they start segments but
|
|
114
|
+
are not provider input). This creates a hash-chain where `prev` links each
|
|
115
|
+
message to its predecessor in the *conversational* history.
|
|
116
|
+
|
|
117
|
+
The index also assigns each message a `fingerprint` (a truncated head/tail
|
|
118
|
+
digest via `Log.truncate_string`) for compact comparison.
|
|
119
|
+
|
|
120
|
+
### Response segments and trace_indices
|
|
121
|
+
|
|
122
|
+
`Chat.trace_indices(indices)` walks a set of message indices and groups them
|
|
123
|
+
into **response segments**. The algorithm:
|
|
124
|
+
|
|
125
|
+
1. Iterate through messages in order.
|
|
126
|
+
2. When a `meta` message is encountered, **close** the current pending segment
|
|
127
|
+
(if any) and **open** a new one, seeded with the parsed metadata.
|
|
128
|
+
3. `user` and `system` messages also close any pending segment.
|
|
129
|
+
4. All other messages (`assistant`, `function_call`, `function_call_output`,
|
|
130
|
+
`tool`, etc.) are appended to the current segment's message list.
|
|
131
|
+
5. At the end, close any remaining pending segment.
|
|
132
|
+
6. Each segment gets `orphan: true` if it has zero covered messages.
|
|
133
|
+
|
|
134
|
+
A `seen` Set of lineage IDs prevents double-counting.
|
|
135
|
+
|
|
136
|
+
### trace_chats
|
|
137
|
+
|
|
138
|
+
```ruby
|
|
139
|
+
def self.trace_chats(chats)
|
|
140
|
+
trace_indices(chats.collect(&:message_index))
|
|
141
|
+
end
|
|
142
|
+
```
|
|
143
|
+
|
|
144
|
+
This takes an array of `Chat` objects, computes `message_index` on each, then
|
|
145
|
+
runs `trace_indices` across all of them. The result is a flat array of segment
|
|
146
|
+
records:
|
|
147
|
+
|
|
148
|
+
```ruby
|
|
149
|
+
{ id: <lineage_id>, meta: <parsed_meta_hash>, messages: [<id>, ...], orphan: true|false }
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
### What trace_chats returns
|
|
153
|
+
|
|
154
|
+
Each entry represents one **response segment** — one model call (or one job
|
|
155
|
+
projection). The `meta` field tells you whether it's a direct inference (has
|
|
156
|
+
`pt`/`ct`/`tt`) or a projection (has `job`). The `messages` array lists the
|
|
157
|
+
lineage IDs of all messages that belong to that segment.
|
|
158
|
+
|
|
159
|
+
**Token accounting from the trace:** To count tokens, filter to entries where
|
|
160
|
+
`meta` has no `:job` key but has `pt`/`ct`/`tt`, then sum those fields. The
|
|
161
|
+
`*_c` and `*_s` counters are checkpoints and must never be summed — they would
|
|
162
|
+
double-count.
|
|
163
|
+
|
|
164
|
+
### Job-based recursive provenance
|
|
165
|
+
|
|
166
|
+
The `Chat.job_chat_files(job, seen)` class method performs recursive traversal
|
|
167
|
+
of the **job dependency graph**:
|
|
168
|
+
|
|
169
|
+
1. Load the job via `Step.load`.
|
|
170
|
+
2. If the job is done and its type is `chat`, add its result path.
|
|
171
|
+
3. Add all `log/**/*.chat` files from the job's `files_dir`.
|
|
172
|
+
4. For each dependency, recurse (using a `seen` Set to prevent cycles and
|
|
173
|
+
duplicate visits).
|
|
174
|
+
5. Return the unique set of all discovered chat file paths.
|
|
175
|
+
|
|
176
|
+
This is the mechanism that `chat.job_chat_files` (instance method) uses to
|
|
177
|
+
discover the full provenance tree below a chat.
|
|
178
|
+
|
|
179
|
+
---
|
|
180
|
+
|
|
181
|
+
## The `prov` Command
|
|
182
|
+
|
|
183
|
+
**File:** `scout_commands/llm/prov`
|
|
184
|
+
|
|
185
|
+
### What it does
|
|
186
|
+
|
|
187
|
+
`prov` prints a hierarchical, indented tree showing the provenance structure
|
|
188
|
+
below a chat file or job path. It walks jobs → agent logs → dependencies
|
|
189
|
+
recursively and prints token totals at each node.
|
|
190
|
+
|
|
191
|
+
### Usage
|
|
192
|
+
|
|
193
|
+
```bash
|
|
194
|
+
scout-ai llm prov <filename>
|
|
195
|
+
```
|
|
196
|
+
|
|
197
|
+
### How it works (internals)
|
|
198
|
+
|
|
199
|
+
The `prov` command **monkey-patches** several methods onto the `Chat` and
|
|
200
|
+
`Step` classes at runtime (these are NOT in the library):
|
|
201
|
+
|
|
202
|
+
- `Chat.provenance(chat_file, prov={})` — recursively walks a chat's jobs,
|
|
203
|
+
their agent chat files, and dependencies, building a hash mapping each chat
|
|
204
|
+
file to the list of agent chat files it references.
|
|
205
|
+
- `Chat.provenance_chat_files(chat)` — flattens the provenance hash into a
|
|
206
|
+
unique list of all chat files.
|
|
207
|
+
- `Chat.tokens(chat)` — loads all provenance chat files, then sums `pt`/`ct`/
|
|
208
|
+
`tt` from direct entries only.
|
|
209
|
+
- `Chat.trace`, `Chat.direct_entries`, `Chat.token_totals`,
|
|
210
|
+
`Chat.print_tokens` — helper methods for trace-based token accounting.
|
|
211
|
+
- `Chat.job_agent_chat_files(job)` — **redefines** the library method.
|
|
212
|
+
- `Step#agent_chats` — new method on Step.
|
|
213
|
+
|
|
214
|
+
These monkey-patches mean `prov` is self-contained but diverges from the
|
|
215
|
+
library's actual API. The `Chat.job_agent_chat_files` in the library already
|
|
216
|
+
exists and works similarly; the `prov` version is a redundant redefinition.
|
|
217
|
+
|
|
218
|
+
### Output format
|
|
219
|
+
|
|
220
|
+
```
|
|
221
|
+
job total=1.2k prompt=800 cont=400 ~/.scout/var/jobs/.../ask
|
|
222
|
+
chat total=500 prompt=300 cont=200 agent.chat
|
|
223
|
+
job total=700 prompt=500 cont=200 ~/.scout/var/jobs/.../subtask
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
- **Yellow `job`** lines show job paths with token totals.
|
|
227
|
+
- **Green `chat`** lines show chat/agent-log files with token totals.
|
|
228
|
+
- Indentation reflects the nesting depth.
|
|
229
|
+
- Token totals are computed by summing `pt`/`ct`/`tt` from direct inference
|
|
230
|
+
metas across all provenance chat files.
|
|
231
|
+
|
|
232
|
+
### Limitations
|
|
233
|
+
|
|
234
|
+
- The `agent.chat` file (the default agent log) is suppressed in the output
|
|
235
|
+
(`unless name == 'agent.chat'`), but its tokens are still counted.
|
|
236
|
+
- The command has a hardcoded fallback filename:
|
|
237
|
+
`~/git/workflows/SC26/chats/network_usecase/3.1.themes` if no argument is
|
|
238
|
+
given.
|
|
239
|
+
- It does not produce machine-readable output (no JSON mode).
|
|
240
|
+
|
|
241
|
+
---
|
|
242
|
+
|
|
243
|
+
## The `info` Command — STATUS ASSESSMENT
|
|
244
|
+
|
|
245
|
+
**File:** `scout_commands/llm/info`
|
|
246
|
+
|
|
247
|
+
### What it does
|
|
248
|
+
|
|
249
|
+
`info` is a **substantially more capable** command than `prov`. It:
|
|
250
|
+
|
|
251
|
+
1. **Discovers the full provenance graph** — root chat, imports, job results,
|
|
252
|
+
job dependencies, and agent logs — using a BFS traversal.
|
|
253
|
+
2. **Deduplicates jobs** by a canonical identity (workflow/task/basename) so
|
|
254
|
+
that `~/.scout` and `~/.rbbt` mirrors of the same job appear once.
|
|
255
|
+
3. **Reports** chats (with role/message counts and job references), jobs (with
|
|
256
|
+
workflow/task, log count, dependency count), and token usage (root-only vs
|
|
257
|
+
all-traced).
|
|
258
|
+
4. **Optional flow mode** (`-f`/`--flow`): prints a compact numbered node/edge
|
|
259
|
+
list with token annotations.
|
|
260
|
+
5. **Optional DOT/plot mode** (`--dot`, `--plot`): generates Graphviz DOT or
|
|
261
|
+
renders SVG/PNG/PDF.
|
|
262
|
+
|
|
263
|
+
### Is `info` outdated? — **NO, it is current and well-maintained**
|
|
264
|
+
|
|
265
|
+
The user suspected `info` might be outdated. After thorough cross-referencing,
|
|
266
|
+
**the `info` command is NOT outdated**. It is in fact the more modern and
|
|
267
|
+
complete implementation. Here is the specific evidence:
|
|
268
|
+
|
|
269
|
+
#### Methods/APIs it calls and their current status
|
|
270
|
+
|
|
271
|
+
| API used by `info` | Defined in | Still exists? |
|
|
272
|
+
|---|---|---|
|
|
273
|
+
| `Chat.load(path)` | `lib/scout/llm/chat/process/meta.rb:86` | ✅ Yes |
|
|
274
|
+
| `Chat.trace_chats(chats)` | `lib/scout/llm/chat/process/meta.rb:195` | ✅ Yes |
|
|
275
|
+
| `Chat.find_file(...)` | `lib/scout/llm/chat/process/files.rb:17` | ✅ Yes |
|
|
276
|
+
| `chat.role_messages(role)` | `lib/scout/llm/chat/annotation.rb` | ✅ Yes |
|
|
277
|
+
| `chat.jobs` / `chat.job_paths` | `lib/scout/llm/chat/process/meta.rb:76` | ✅ Yes |
|
|
278
|
+
| `job.dependencies` | Scout `Step` API | ✅ Yes |
|
|
279
|
+
| `job.file('log')` | Scout `Step` API | ✅ Yes |
|
|
280
|
+
| `job.info[:workflow]`, `job.info[:task_name]` | Scout `Step` API | ✅ Yes |
|
|
281
|
+
| `job.done?`, `job.type` | Scout `Step` API | ✅ Yes |
|
|
282
|
+
| `Step.load(path)` | Scout API | ✅ Yes |
|
|
283
|
+
|
|
284
|
+
#### Annotation fields it reads
|
|
285
|
+
|
|
286
|
+
The `info` command reads `pt`, `ct`, `tt` from parsed meta messages via
|
|
287
|
+
`Chat.trace_chats` → `trace_indices` → per-entry `:meta` hash. These fields are
|
|
288
|
+
written by `Backend::Default#update_meta` (line 420 of `backends/default.rb`)
|
|
289
|
+
and are fully current.
|
|
290
|
+
|
|
291
|
+
It correctly filters to **direct entries only** (entries whose meta has no
|
|
292
|
+
`:job` key but has `pt`/`ct`/`tt`), matching the exact pattern used by
|
|
293
|
+
`ChatAnalyst::Session#token_entries`.
|
|
294
|
+
|
|
295
|
+
#### What `info` does that `prov` does not
|
|
296
|
+
|
|
297
|
+
| Feature | `info` | `prov` |
|
|
298
|
+
|---|---|---|
|
|
299
|
+
| Import discovery | ✅ (`import`/`continue`/`last` roles) | ❌ |
|
|
300
|
+
| Job deduplication | ✅ (canonical identity by workflow/task/basename) | ❌ |
|
|
301
|
+
| Dependency graph | ✅ | ✅ (via Chat.provenance) |
|
|
302
|
+
| Flow visualization | ✅ (text flow + Graphviz DOT/SVG/PNG/PDF) | ❌ |
|
|
303
|
+
| Warnings | ✅ (records load failures) | ❌ |
|
|
304
|
+
| Token accounting | ✅ (`trace_chats` + direct_entries) | ✅ (same approach, but monkey-patched) |
|
|
305
|
+
| Uses library API directly | ✅ (no monkey-patching) | ❌ (redefines Chat methods) |
|
|
306
|
+
|
|
307
|
+
#### Assessment summary
|
|
308
|
+
|
|
309
|
+
- **`info` is the current, recommended command** for provenance inspection.
|
|
310
|
+
- **`prov` is older** and relies on runtime monkey-patches rather than the
|
|
311
|
+
library API. It still works but is superseded by `info` in functionality.
|
|
312
|
+
- Neither command is "broken" — both will execute successfully.
|
|
313
|
+
- If anything is "outdated," it is `prov` (monkey-patches, no import
|
|
314
|
+
discovery, no flow/DOT output), not `info`.
|
|
315
|
+
|
|
316
|
+
---
|
|
317
|
+
|
|
318
|
+
## ChatAnalyst Provenance Capabilities
|
|
319
|
+
|
|
320
|
+
**Location:** `~/git/workflows/SC26/Agent/ChatAnalyst/`
|
|
321
|
+
|
|
322
|
+
### What the ChatAnalyst agent does
|
|
323
|
+
|
|
324
|
+
ChatAnalyst is a Scout-AI agent specialized in **inspecting work sessions** to
|
|
325
|
+
find usage patterns, tool-calling issues, barriers, and areas for improving
|
|
326
|
+
the harness, agent instructions, or tooling.
|
|
327
|
+
|
|
328
|
+
Its system prompt (`start_chat`) directs it to:
|
|
329
|
+
- Follow conversations across different chats and ask jobs.
|
|
330
|
+
- Inspect persisted chat sessions for patterns and issues.
|
|
331
|
+
- Use the ChatAnalyst workflow tooling (README.md) for structured analysis.
|
|
332
|
+
|
|
333
|
+
### The Session class (workflow.rb core)
|
|
334
|
+
|
|
335
|
+
The `workflow.rb` defines a `ChatAnalyst::Session` class that performs
|
|
336
|
+
BFS-based discovery identical in spirit to `info`'s `LLMInfoReport`:
|
|
337
|
+
|
|
338
|
+
1. **`resolve_chat(input)`** — tries multiple candidate paths (plain,
|
|
339
|
+
`.chat` extension, expanded, `Scout.chats[]` lookup).
|
|
340
|
+
|
|
341
|
+
2. **`discover_chat(path)`** — for each chat:
|
|
342
|
+
- Records it in `@chats`.
|
|
343
|
+
- Follows `import`/`continue`/`last` roles to discover imported chats
|
|
344
|
+
(adding `import` edges).
|
|
345
|
+
- Follows `meta job=...` references to discover producer jobs (adding
|
|
346
|
+
`result` edges).
|
|
347
|
+
|
|
348
|
+
3. **`discover_job(reference)`** — for each job:
|
|
349
|
+
- Records it in `@jobs`.
|
|
350
|
+
- Follows dependencies recursively (adding `dependency` edges).
|
|
351
|
+
- If the job is done and type is `chat`, discovers the result chat (adding
|
|
352
|
+
it to the chat queue).
|
|
353
|
+
- Scans `log/**/*.chat` for agent conversation logs (adding `log` edges and
|
|
354
|
+
feeding them back into `discover_chat`).
|
|
355
|
+
|
|
356
|
+
4. **Token accounting** — `token_entries` filters the trace to direct inference
|
|
357
|
+
segments (no `:job` key, has `pt`/`ct`/`tt`). `token_totals` sums them.
|
|
358
|
+
|
|
359
|
+
### Tasks exposed by the ChatAnalyst workflow
|
|
360
|
+
|
|
361
|
+
| Task | Type | Purpose |
|
|
362
|
+
|---|---|---|
|
|
363
|
+
| `message_index` | JSON | Compact per-message index across the full session tree. Each entry has an ID (`file#index`), lineage ID, previous lineage ID, role, fingerprint, and parsed meta. Supports optional `role` filter. |
|
|
364
|
+
| `message_content` | JSON | Retrieves full untruncated content for specific message IDs. Two-phase inspection: scan index → drill into interesting messages. |
|
|
365
|
+
| `chat_overview` | JSON | Structural graph overview: per-file summaries (roles, messages, jobs, tool calls), per-job summaries (workflow/task, dependencies, logs), typed edges, aggregate totals, warnings. |
|
|
366
|
+
| `chat_tool_calls` | JSON | Pairs `function_call`/`mcp_call` with `function_call_output` by call ID. Reports tool name, call ID, output position, success/failure, exception/exit status. Summary: totals, successes, failures, breakdown by tool. |
|
|
367
|
+
| `chat_tokens` | JSON | Per-file and aggregate token accounting using direct `pt`/`ct`/`tt` metas only. Includes `note` explaining methodology. |
|
|
368
|
+
| `chat_agents` | JSON | Detects `ask` and `hand_off_to_*` calls. Reports interactions with source file, call ID, output position, success state. |
|
|
369
|
+
| `chat_report` | JSON | Combined snapshot: session size, job count, aggregate tokens, trace records, first 20 tool calls, all failures, warnings. |
|
|
370
|
+
|
|
371
|
+
### Tooling file
|
|
372
|
+
|
|
373
|
+
The `tooling` file in the ChatAnalyst agent directory is a **chat session
|
|
374
|
+
log** (not a tooling declaration file). It contains the conversation in which
|
|
375
|
+
the ChatAnalyst workflow was originally designed, including a detailed
|
|
376
|
+
critique and README update session. It is not loaded as tooling — it is a
|
|
377
|
+
record of the development process.
|
|
378
|
+
|
|
379
|
+
### How ChatAnalyst compares to `info`
|
|
380
|
+
|
|
381
|
+
| Aspect | `info` CLI | ChatAnalyst workflow |
|
|
382
|
+
|---|---|---|
|
|
383
|
+
| Consumer | Human (terminal output, DOT/SVG) | Agent (JSON tasks) |
|
|
384
|
+
| Discovery | Identical BFS (imports, jobs, deps, logs) | Identical BFS |
|
|
385
|
+
| Token accounting | `trace_chats` + direct entries | `trace_chats` + direct entries |
|
|
386
|
+
| Job deduplication | ✅ (canonical identity) | ❌ (simple expand_path check) |
|
|
387
|
+
| Tool-call analysis | ❌ | ✅ |
|
|
388
|
+
| Agent-interaction analysis | ❌ | ✅ |
|
|
389
|
+
| Output | Text / Graphviz | JSON |
|
|
390
|
+
| Graphviz rendering | ✅ | ❌ |
|
|
391
|
+
|
|
392
|
+
---
|
|
393
|
+
|
|
394
|
+
## Key Design Notes
|
|
395
|
+
|
|
396
|
+
### The provenance tree metaphor
|
|
397
|
+
|
|
398
|
+
A Scout-AI session is not a single flat conversation. It is a **tree** (or DAG)
|
|
399
|
+
of conversations connected by:
|
|
400
|
+
|
|
401
|
+
1. **Import edges** — `import:`/`continue:`/`last:` roles pull in previous
|
|
402
|
+
chat history, creating a horizontal chain of related conversations.
|
|
403
|
+
|
|
404
|
+
2. **Result edges** — `meta: job=<path>` markers indicate that a response
|
|
405
|
+
segment was produced by a Scout workflow job (typically an `ask` task).
|
|
406
|
+
The actual inference happened inside that job's agent logs.
|
|
407
|
+
|
|
408
|
+
3. **Dependency edges** — Scout workflow jobs have dependencies (other jobs
|
|
409
|
+
they consume). Each dependency may itself be a chat-producing job with its
|
|
410
|
+
own agent logs.
|
|
411
|
+
|
|
412
|
+
4. **Log edges** — Each ask-job has a `log/` directory containing agent chat
|
|
413
|
+
files. The primary `log/agent.chat` is the agent's full conversation.
|
|
414
|
+
Socialized projections may appear under `log/chats/<AgentName>/`.
|
|
415
|
+
|
|
416
|
+
5. **Call edges** (semantic, not structural) — Tool calls to `ask` or
|
|
417
|
+
`hand_off_to_*` represent agent-to-agent delegation. These are detected by
|
|
418
|
+
scanning tool-call messages, not by following meta references.
|
|
419
|
+
|
|
420
|
+
### How multi-agent inference creates provenance chains
|
|
421
|
+
|
|
422
|
+
When an orchestrator agent (e.g., `AGI`) dispatches work to a specialist agent
|
|
423
|
+
(e.g., `Worker`) via `ask`:
|
|
424
|
+
|
|
425
|
+
1. The orchestrator's chat contains a `function_call` with tool name `ask` and
|
|
426
|
+
arguments naming the target agent.
|
|
427
|
+
2. The `ask` call triggers a Scout workflow job (e.g.,
|
|
428
|
+
`Agent/Worker/ask`).
|
|
429
|
+
3. That job runs the target agent, producing `log/agent.chat` with the full
|
|
430
|
+
agent conversation and its own `meta:` entries with direct token counts.
|
|
431
|
+
4. The job result is projected back into the orchestrator's chat as a
|
|
432
|
+
`meta: job=<path>` marker followed by the response messages.
|
|
433
|
+
5. The orchestrator's chat thus has a **zero-token projection segment** — the
|
|
434
|
+
real tokens are in the job's `agent.chat`.
|
|
435
|
+
|
|
436
|
+
This means token accounting requires **recursive traversal**: start at the root
|
|
437
|
+
chat, follow every `meta: job=...` to its job, read the job's `agent.chat`
|
|
438
|
+
logs, follow the job's dependencies, and sum only direct `pt`/`ct`/`tt`
|
|
439
|
+
entries. Cumulative (`*_c`) and session (`*_s`) counters must never be summed
|
|
440
|
+
because they would double-count across the import/projection chain.
|
|
441
|
+
|
|
442
|
+
### Socialized chat projections
|
|
443
|
+
|
|
444
|
+
When an agent dispatches to another agent with a named `conversation`, the
|
|
445
|
+
specialist interaction may be persisted as a socialized chat at:
|
|
446
|
+
```
|
|
447
|
+
<caller_job>.files/log/chats/<AgentName>/<conversation>.chat
|
|
448
|
+
```
|
|
449
|
+
These files contain the prompt, propagated `option:` lines, a `meta: job=...`
|
|
450
|
+
marker, and the response — but **zero direct inference tokens**. The real
|
|
451
|
+
model calls are found by following the `meta: job=...` reference into the
|
|
452
|
+
specialist's ask-job and its `agent.chat` log.
|
|
453
|
+
|
|
454
|
+
### Why `info` is preferred over `prov`
|
|
455
|
+
|
|
456
|
+
- `info` uses the library API directly (no monkey-patching).
|
|
457
|
+
- `info` discovers imports (prov does not).
|
|
458
|
+
- `info` deduplicates mirrored jobs between `~/.scout` and `~/.rbbt`.
|
|
459
|
+
- `info` supports flow visualization (text + Graphviz).
|
|
460
|
+
- `info` records and reports warnings for load failures.
|
|
461
|
+
- `prov` remains functional but is architecturally older and less complete.
|
|
462
|
+
|
|
463
|
+
### Design consistency between `info` and ChatAnalyst
|
|
464
|
+
|
|
465
|
+
Both `info` and ChatAnalyst use the same `Chat.trace_chats` API for token
|
|
466
|
+
accounting and the same BFS pattern for provenance discovery. The key
|
|
467
|
+
difference is that `info` adds job deduplication (canonical identity), while
|
|
468
|
+
ChatAnalyst adds tool-call and agent-interaction analysis. They are
|
|
469
|
+
complementary tools for the same provenance data model.
|